- Generate Structured Instruction
Description
Translates a user's text-based edit instruction and source image/mask into a detailed, machine-readable structured edit instruction in JSON format.
This endpoint uses the state-of-the-art Gemini 2.5 Flash VLM bridge to understand the edit context. It only returns the JSON string and does not generate an image.
Context-Aware Masking
When a mask is provided, the VLM analyzes the specific region of interest in relation to the rest of the image. It generates a structured_instruction tailored specifically for that area (e.g., ensuring lighting and perspective match the unmasked background), ensuring seamless integration when the edit is applied.
Why use this endpoint?
- Decoupling: Decouples the "intent translation" step from the "image editing" step, giving you maximum flexibility.
- Control & Auditability: Allows for a "human-in-the-loop" to inspect, programmatically edit, or version the JSON before generating an image (e.g., for a custom UI).
- Consistency & Automation: Generate one
structured_instructionand pass it to/v2/image/editmultiple times to create consistent, auditable variations. - Hybrid Deployment: Use Bria's state-of-the-art VLM bridge via API while self-hosting the open-source FIBO image model on your own private cloud.
The resulting structured_instruction can be used as input for the /v2/image/edit endpoint.
Input Combination Rules The request body must use exactly one of the following combinations:
- Global Instruction:
images+instruction - Masked Instruction:
images+mask+instruction
API Access
You can register and access the API Token through Bria's platform by clicking here.
Required. Text-based edit instruction (e.g., "make the sky blue", "add a cat"). This parameter serves as the text prompt.
The source/reference image(s) to be edited. Publicly available URL or Base64-encoded. Accepted formats: JPEG, JPG, PNG, WEBP. Accepts 1 to 4 images.
Publicly available URL or Base64-encoded mask image (black and white). Black areas will be preserved, white areas will be edited. If omitted, the edit applies to the entire image. The input image and the input mask must be of the same size. Only supported when images contains exactly one item.
Optional. Seed for deterministic generation. If omitted, a random seed is generated and used.
Specifies the response mode. Optional.
- When
false(default), the request is processed asynchronously: the API immediately returns a status URL to track progress. - When
true, the request is processed synchronously: the API hold the connection open until the proccess is complete and then returns the final result in the response.
Optional URL for receiving the result via webhook when the async job completes. See Webhooks.
If true, returns a warning for potential IP content in the instruction. Optional.
If true, returns 422 on instruction moderation failure. Optional.
- Generate Instruction from Image & Text
- Generate Instruction from Masked Image & Text
curl -i -X POST \
https://engine.prod.bria-api.com/v2/structured_instruction/generate \
-H 'Content-Type: application/json' \
-H 'api_token: string' \
-d '{
"images": [
"https://bria-datasets.s3.us-east-1.amazonaws.com/api_doc/fibo-edit/42082.jpg"
],
"instruction": "create a detailed realistic photo, with contemporary color scheme, and balanced exposure photo roughly based on this sketch"
}'{ "result": { "seed": 0, "structured_instruction": "string" }, "request_id": "string", "warning": "string" }