Developer Docs/Image to Video API

Image to Video API

Use the A2E Image to Video API to convert a still image into an AI-generated video with prompts, duration controls, and result management endpoints.

Overview

Convert a still image into an AI-animated video. Control motion intensity, duration, and aspect ratio to produce smooth, high-quality video clips from a single reference image.

Primary Endpoint

POST/api/v1/userImage2Video/start

Convert an image to video using AI generation with custom prompts. ### Request - Send a valid bearer token. The server evaluates the operation in the authenticated caller's access context. - Send an `application/json` body when using the optional controls documented in the request schema. - API-token callers may include `webhook_url` and `webhook_token` for best-effort terminal-state notifications; ordinary JWT/cookie calls ignore these fields. ### Behavior - This is an asynchronous operation: a successful submission creates a task and returns before processing finishes. - Persist the returned task identifier and use the corresponding detail or list operation to observe progress. - Treat the detail endpoint as the source of truth even when webhook delivery is enabled. ### Response - A `200` response confirms task acceptance; it does not by itself mean media generation has completed. - Retain the returned identifier and wait for a documented terminal status before using output URLs. - JSON object responses, including error responses, normally carry a top-level `trace_id` string for this request; include it when contacting support. It is not a task identifier. - Do not infer undocumented fields or statuses; clients should tolerate additional response properties. ### Errors - `400` — Bad Request - Invalid parameters. - `401` — Unauthorized - Invalid or missing JWT token. ### Related Operations - `GET /api/v1/userImage2Video/allRecords` — List image to video tasks. - `GET /api/v1/userImage2Video/{_id}` — Get image to video task detail. - `DELETE /api/v1/userImage2Video/{_id}` — Delete an image to video task. Authentication: set header Authorization: Bearer <token> (supports user JWT or sk_ API token).

Request Parameters

NameTypeRequiredDescription
namestringNoName of the image to video task
image_urlstringNoURL of the source image. Required for image-to-video; omit for text-to-video and reference-to-video.
promptstringNoGeneration prompt
negative_promptstringNoNegative prompt to avoid unwanted features
lorastringNoLora ID. If provided, prompt/negative_prompt can be omitted (prompt will be overwritten by lora preset on backend).
generation_modeenum: image-to-video | text-to-video | reference-to-videoNoGeneration mode. Text and reference modes require a2e-v2 or a2e-v2-flash and a nonempty prompt up to 2000 characters. Reference mode accepts video_time of 3, 4, 5, 10 or 15; text mode accepts 5, 10 or 15.; default: "image-to-video"
generation_paramsobjectNoText/reference parameters. Omit for image-to-video. Reference mode requires at least one asset; text mode accepts none.
generation_params.aspect_ratioenum: 1:1 | 2:3 | 3:2 | 3:4 | 4:3 | 9:16 | 16:9 | 21:9Nodefault: "16:9"
generation_params.reference_assetsarray<object>NoAt most 9 images, 3 videos and 3 audio files. Video and audio clips must each last 2–15 seconds, with at most 15 seconds total per type.; maxItems: 12
generation_params.reference_assets[].asset_idstringYesUnique reference identifier, such as image_1, video_1 or audio_1.
generation_params.reference_assets[].typeenum: image | video | audioYes-
generation_params.reference_assets[].urlstringYes-
generation_params.reference_assets[].duration_secondsnumberNoRequired for audio. Video duration is measured by the server.; minimum: 2; maximum: 15
model_typeenum: GENERAL | FLF2VNoModel type for generation; default: "GENERAL"
model_versionenum: a2e | a2e-v2 | a2e-v2-flashNoAlgorithm model version. When omitted, Ultra users default to a2e-v2 and other roles default to a2e. a2e-v2 costs more per second; a2e-v2-flash uses a2e pricing. V2 models are not covered by plan waivers.
auto_add_audiobooleanNoWhether to add audio. Defaults to false for a2e (paid ThinkSound post-processing) and true for a2e-v2/a2e-v2-flash (native audio with no ThinkSound surcharge).
end_image_urlstringNoEnd image URL (required for FLF2V model)
extend_promptbooleanNoWhether to extend the prompt automatically; default: true
number_of_imagesintegerNoNumber of videos to generate at once; minimum: 1; maximum: 8; default: 1
video_timeintegerNoRequested video time in integer seconds; billing uses this value. Image-to-video and first-last-frame (FLF2V) support 3-20 seconds; video extension supports 5-20 seconds. Reference mode supports 3, 4, 5, 10 or 15 seconds; text mode supports 5, 10 or 15. V1 3/4-second videos use 49/65 frames at 16fps; V1 durations above 5 seconds retain 5-second segment rounding. V2 output duration rounds up to the next supported 17k+5 frame count at 24fps.; minimum: 3; maximum: 20; default: 5
video_lengthintegerNo(Deprecated) Video length in frames. Prefer video_time. Backend converts frames to seconds internally.; minimum: 1; maximum: 1000
skip_face_enhancebooleanNoWhether to skip face similarity enhancement. Defaults to false (enhancing face similarity).; default: false
mask_facebooleanNoWeb NSFW preflight result. Set true only after the user confirms masking a detected human face.
minor_suspected_skipbooleanNoAccepted for backward compatibility only. This endpoint has no backend CSAM detection; moderation comes from the algorithm upstream and this flag is not forwarded to it, so setting it does not bypass an upstream 1004.; default: false
webhook_urlstringNoHTTPS URL to receive task.completed / task.failed notifications. Best-effort delivery, single attempt, no retries; clients should treat the detail API as the source of truth.; maxLength: 2048
webhook_tokenstringNoOptional plaintext token returned in the X-A2e-Webhook-Token header so receivers can verify the request originated from a2e.; maxLength: 256
Request schema and conditional rules
{
  "allOf": [
    {
      "type": "object",
      "properties": {
        "name": {
          "type": "string",
          "description": "Name of the image to video task",
          "example": "My Image Animation"
        },
        "image_url": {
          "type": "string",
          "description": "URL of the source image. Required for image-to-video; omit for text-to-video and reference-to-video.",
          "example": "https://example.com/image.jpg"
        },
        "prompt": {
          "type": "string",
          "description": "Generation prompt",
          "example": "Make the person in the image wave their hand"
        },
        "negative_prompt": {
          "type": "string",
          "description": "Negative prompt to avoid unwanted features",
          "example": "blurry, distorted, static"
        },
        "lora": {
          "type": "string",
          "description": "Lora ID. If provided, prompt/negative_prompt can be omitted (prompt will be overwritten by lora preset on backend)."
        },
        "generation_mode": {
          "type": "string",
          "enum": [
            "image-to-video",
            "text-to-video",
            "reference-to-video"
          ],
          "default": "image-to-video",
          "description": "Generation mode. Text and reference modes require a2e-v2 or a2e-v2-flash and a nonempty prompt up to 2000 characters. Reference mode accepts video_time of 3, 4, 5, 10 or 15; text mode accepts 5, 10 or 15."
        },
        "generation_params": {
          "type": "object",
          "description": "Text/reference parameters. Omit for image-to-video. Reference mode requires at least one asset; text mode accepts none.",
          "properties": {
            "aspect_ratio": {
              "type": "string",
              "enum": [
                "1:1",
                "2:3",
                "3:2",
                "3:4",
                "4:3",
                "9:16",
                "16:9",
                "21:9"
              ],
              "default": "16:9"
            },
            "reference_assets": {
              "type": "array",
              "maxItems": 12,
              "description": "At most 9 images, 3 videos and 3 audio files. Video and audio clips must each last 2–15 seconds, with at most 15 seconds total per type.",
              "items": {
                "type": "object",
                "required": [
                  "asset_id",
                  "type",
                  "url"
                ],
                "properties": {
                  "asset_id": {
                    "type": "string",
                    "description": "Unique reference identifier, such as image_1, video_1 or audio_1."
                  },
                  "type": {
                    "type": "string",
                    "enum": [
                      "image",
                      "video",
                      "audio"
                    ]
                  },
                  "url": {
                    "type": "string",
                    "format": "uri"
                  },
                  "duration_seconds": {
                    "type": "number",
                    "minimum": 2,
                    "maximum": 15,
                    "description": "Required for audio. Video duration is measured by the server."
                  }
                }
              }
            }
          }
        },
        "model_type": {
          "type": "string",
          "enum": [
            "GENERAL",
            "FLF2V"
          ],
          "default": "GENERAL",
          "description": "Model type for generation"
        },
        "model_version": {
          "type": "string",
          "enum": [
            "a2e",
            "a2e-v2",
            "a2e-v2-flash"
          ],
          "description": "Algorithm model version. When omitted, Ultra users default to a2e-v2 and other roles default to a2e. a2e-v2 costs more per second; a2e-v2-flash uses a2e pricing. V2 models are not covered by plan waivers."
        },
        "auto_add_audio": {
          "type": "boolean",
          "description": "Whether to add audio. Defaults to false for a2e (paid ThinkSound post-processing) and true for a2e-v2/a2e-v2-flash (native audio with no ThinkSound surcharge)."
        },
        "end_image_url": {
          "type": "string",
          "description": "End image URL (required for FLF2V model)",
          "example": "https://example.com/end_image.jpg"
        },
        "extend_prompt": {
          "type": "boolean",
          "default": true,
          "description": "Whether to extend the prompt automatically"
        },
        "number_of_images": {
          "type": "integer",
          "minimum": 1,
          "maximum": 8,
          "default": 1,
          "description": "Number of videos to generate at once",
          "example": 1
        },
        "video_time": {
          "type": "integer",
          "minimum": 3,
          "maximum": 20,
          "default": 5,
          "description": "Requested video time in integer seconds; billing uses this value. Image-to-video and first-last-frame (FLF2V) support 3-20 seconds; video extension supports 5-20 seconds. Reference mode supports 3, 4, 5, 10 or 15 seconds; text mode supports 5, 10 or 15. V1 3/4-second videos use 49/65 frames at 16fps; V1 durations above 5 seconds retain 5-second segment rounding. V2 output duration rounds up to the next supported 17k+5 frame count at 24fps.",
          "example": 5
        },
        "video_length": {
          "type": "integer",
          "minimum": 1,
          "maximum": 1000,
          "description": "(Deprecated) Video length in frames. Prefer video_time. Backend converts frames to seconds internally.",
          "example": 81
        },
        "skip_face_enhance": {
          "type": "boolean",
          "default": false,
          "description": "Whether to skip face similarity enhancement. Defaults to false (enhancing face similarity)."
        },
        "mask_face": {
          "type": "boolean",
          "description": "Web NSFW preflight result. Set true only after the user confirms masking a detected human face."
        },
        "minor_suspected_skip": {
          "type": "boolean",
          "default": false,
          "description": "Accepted for backward compatibility only. This endpoint has no backend CSAM detection; moderation comes from the algorithm upstream and this flag is not forwarded to it, so setting it does not bypass an upstream 1004."
        }
      },
      "required": []
    },
    {
      "$ref": "#/components/schemas/WebhookInput"
    }
  ]
}

Response Fields

code: integer
data: object
data._id: string
data.name: string
data.image_url: string
data.current_status: string
data.result_url: string
data.cover_url: string
data.mask_face: boolean
Whether clients should display the privacy-processed input image as the task cover
data.model_type: string
data.end_image_url: string
data.extend_prompt: boolean
data.lora: string
data.video_time: number
data.video_length: number
data.coins: number
data.remainingDays: number
data.expirationDate: string
data.isExpired: boolean
data.expirationDays: number
data.total_videos: number
Present only when number_of_images > 1
data.all_records: array<object>
Present only when number_of_images > 1
trace_id: string
Trace ID of this HTTP request. Include it when contacting support about this request. It is generated per request and is not a task identifier; use the returned task `_id` to query results.

Request Example

curl -X POST "https://www.a2e.com.cn/api/v1/userImage2Video/start" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
  "name": "My Image Animation",
  "image_url": "https://example.com/image.jpg",
  "prompt": "Make the person in the image wave their hand",
  "negative_prompt": "blurry, distorted, static",
  "model_type": "GENERAL",
  "extend_prompt": true,
  "number_of_images": 1,
  "video_time": 5,
  "skip_face_enhance": false,
  "mask_face": false
}'

Related Endpoints

Responses

200

Image to video task started successfully

400

Bad Request - Invalid parameters

401

Unauthorized - Invalid or missing bearer token

Image to Video API Documentation