Skip to main content
POST
Analyze multimodal input with Gemini
Use the Responses endpoint with gemini-3.8-flash, gemini-3.7-flash, or gemini-3.5-flash-lite.

Video understanding request

The media input can be a public or signed HTTPS URL, or a Base64 data URL. Raw gs:// URIs are not accepted; use a signed HTTPS URL instead.

Response

Parameters

Video parts accept fps in the range (0, 24], startOffset, endOffset, and mediaResolution (low, medium, or high). Google defaults video sampling to 1 FPS when fps is omitted. Put the text instruction after the video part for the most reliable video analysis. Gemini 3.5 Flash-Lite rejects a request whose final non-system message uses the assistant role. Use minimal reasoning for extraction and other straightforward, latency-sensitive tasks. See the Gemini 3.8 Flash, Gemini 3.7 Flash and Gemini 3.5 Flash-Lite model pages for supported modalities. Gemini 3.8 Flash defaults to medium reasoning and rejects minimal, unsupported reasoning levels, and thinkingBudget overrides. Use low, medium, or high.

Function calling and built-in tools

Function calling, Code execution, and Computer use are supported for all three Gemini models. See Function calling and Gemini tools for complete requests, multi-turn result submission, thought signature preservation, streaming behavior, and billing.

Streaming Responses

Set stream: true to receive typed Responses events as the model generates output:
The event sequence includes response.created, response.in_progress, response.output_item.added, content or function-argument deltas, corresponding done events, and a terminal response.completed, response.incomplete, or response.failed. Sequence numbers order events; item IDs and output indexes remain stable. The final response contains the full output and usage. For Function Call, consume response.function_call_arguments.delta and response.function_call_arguments.done. Gemini currently emits complete function arguments in a single delta per call. Computer use actions follow the same format. Code execution and media artifacts remain available in the output items’ providerMetadata.google.parts. Retain all final output items for the next turn. Wallet settlement completes before the terminal response event is emitted. If streaming or settlement fails, the endpoint emits event: error and closes without a successful terminal event. Missing provider usage is retained for reconciliation. Responses uses typed terminal events rather than Chat Completions’ [DONE] marker. background: true remains unsupported. This is a live HTTP stream subject to the Vercel function time limit, not a persisted background task or a resumable stream.

Authorizations

Authorization
string
header
required

A workspace API key beginning with gengen_live_.

Body

application/json
model
enum<string>
required

Gemini model used for the response.

Available options:
gemini-3.8-flash,
gemini-3.7-flash,
gemini-3.5-flash-lite
input
required

Text or multimodal conversation input.

stream
boolean
default:false

Return typed Responses SSE events, including function argument deltas and terminal response usage.

background
boolean
default:false

Background execution is not supported.

instructions
string

High-level instructions sent as the Gemini system instruction.

reasoning_effort
enum<string>

Thinking level. minimal is supported by Gemini 3.5 Flash-Lite but not Gemini 3.8 Flash or Gemini 3.7 Flash. Gemini 3.8 Flash defaults to medium and rejects thinking budgets.

Available options:
minimal,
low,
medium,
high
temperature
number

Sampling temperature. Custom values are ignored by Gemini 3.5 Flash-Lite.

top_p
number

Nucleus-sampling probability threshold. Custom values are ignored by Gemini 3.5 Flash-Lite.

Required range: 0 <= x <= 1
top_k
number

Limits sampling to the most likely K tokens. Custom values are ignored by Gemini 3.5 Flash-Lite.

Required range: x >= 1
max_output_tokens
integer

Maximum output tokens. max_tokens and max_completion_tokens are also accepted.

Required range: x >= 1
max_tokens
integer
Required range: x >= 1
stop

One or more sequences that stop generation.

seed
integer

Best-effort sampling seed.

text
object

Structured-output settings for supported Gemini models.

tools
object[]

Function declarations. Chat uses nested function objects; Responses uses flat function declarations. Built-in tools belong in providerOptions.google.tools.

tool_choice

auto, none, required, or a named function object. The function name must be declared in tools.

Available options:
auto,
none,
required
parallel_tool_calls
boolean

Omit or use true. Gemini can return multiple calls; false is rejected.

providerOptions
object

Response

An OpenAI-compatible Responses object or typed SSE stream.

id
string
object
string
Allowed value: "response"
created_at
integer
status
enum<string>
Available options:
completed,
incomplete,
failed
model
string
output
object[]
output_text
string
usage
object