> ## Documentation Index
> Fetch the complete documentation index at: https://docs.gengen.farm/llms.txt
> Use this file to discover all available pages before exploring further.

# Multimodal responses with Gemini

> Analyze text, image, video, and audio inputs with Gemini 3.8 Flash, Gemini 3.7 Flash, or Gemini 3.5 Flash-Lite through the GENGEN Responses API.

Use the Responses endpoint with `gemini-3.8-flash`, `gemini-3.7-flash`, or `gemini-3.5-flash-lite`.

| Property | Value |
| - | - |
| Method | `POST` |
| Endpoint | `/api/gengen/v1/responses` |
| Models | `gemini-3.8-flash`, `gemini-3.7-flash`, `gemini-3.5-flash-lite` |
| Processing | Synchronous or SSE streaming |

## Video understanding request

The media input can be a public or signed HTTPS URL, or a Base64 data URL. Raw `gs://` URIs are not accepted; use a signed HTTPS URL instead.

```bash theme={"dark"}
curl --request POST \
  --url https://gengen.farm/api/gengen/v1/responses \
  --header 'Authorization: Bearer gengen_live_xxxxxxxxxxxxxxxx' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "gemini-3.8-flash",
    "input": [
      {
        "role": "user",
        "content": [
          {
            "type": "input_video",
            "video_url": "https://example.com/input.mp4"
          },
          {
            "type": "input_text",
            "text": "Describe the sequence of actions in this video."
          }
        ]
      }
    ],
    "reasoning_effort": "medium"
  }'
```

## Response

```json theme={"dark"}
{
  "id": "google-response-id",
  "object": "response",
  "status": "completed",
  "model": "gemini-3.8-flash",
  "output": [
    {
      "type": "message",
      "status": "completed",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "A person enters the greenhouse, waters the plants, and checks the temperature display."
        }
      ]
    }
  ],
  "output_text": "A person enters the greenhouse, waters the plants, and checks the temperature display."
}
```

## Parameters

| Parameter | Type | Description |
| - | - | - |
| `model` | string | `gemini-3.8-flash`, `gemini-3.7-flash`, or `gemini-3.5-flash-lite`. |
| `input` | string or array | Required text or multimodal input. Content parts may be `input_text`, `input_image`, `input_video`, or `input_audio`. |
| `stream` | boolean | Set to `true` for Responses SSE events. Defaults to `false`. |
| `instructions` | string | High-level instructions sent as the Gemini system instruction. |
| `reasoning_effort` | string | `low`, `medium`, or `high` for Gemini 3.8 Flash and Gemini 3.7 Flash; Gemini 3.5 Flash-Lite also supports and defaults to `minimal`. |
| `temperature`, `top_p`, `top_k` | number | Sampling controls. Gemini 3.5 Flash-Lite ignores custom values. |
| `max_tokens` | integer | Maximum generated tokens. |
| `stop` | string or string\[] | One or more stop sequences. |
| `seed` | integer | Best-effort sampling seed. |
| `text.format` | object | `json_object` or `json_schema` structured output. |
| `providerOptions.google` | object | Advanced Google `generationConfig`, `safetySettings`, `tools`, and `toolConfig`. |

Video parts accept `fps` in the range `(0, 24]`, `startOffset`, `endOffset`, and `mediaResolution` (`low`, `medium`, or `high`). Google defaults video sampling to 1 FPS when `fps` is omitted. Put the text instruction after the video part for the most reliable video analysis.

Gemini 3.5 Flash-Lite rejects a request whose final non-system message uses the `assistant` role. Use `minimal` reasoning for extraction and other straightforward, latency-sensitive tasks.

See the [Gemini 3.8 Flash](/models/gemini-3.8-flash), [Gemini 3.7 Flash](/models/gemini-3.7-flash) and [Gemini 3.5 Flash-Lite](/models/gemini-3.5-flash-lite) model pages for supported modalities.

Gemini 3.8 Flash defaults to `medium` reasoning and rejects `minimal`, unsupported reasoning levels, and `thinkingBudget` overrides. Use `low`, `medium`, or `high`.

## Function calling and built-in tools

Function calling, Code execution, and Computer use are supported for all three Gemini models.
See [Function calling and Gemini tools](/tool-calling) for complete requests, multi-turn
result submission, thought signature preservation, streaming behavior, and billing.

## Streaming Responses

Set `stream: true` to receive typed Responses events as the model generates output:

```bash theme={"dark"}
curl --no-buffer https://gengen.farm/api/gengen/v1/responses \
  --header "Authorization: Bearer $GENGEN_API_KEY" \
  --header 'Content-Type: application/json' \
  --data '{"model":"gemini-3.8-flash","input":"Explain idempotency briefly.","stream":true}'
```

```text theme={"dark"}
event: response.created
data: {"type":"response.created","sequence_number":0,"response":{"id":"google-response-id","object":"response","status":"in_progress","output":[]}}

event: response.output_text.delta
data: {"type":"response.output_text.delta","sequence_number":4,"item_id":"msg_google-response-id_0","output_index":0,"content_index":0,"delta":"Idempotency means"}
```

The event sequence includes `response.created`, `response.in_progress`,
`response.output_item.added`, content or function-argument deltas, corresponding
`done` events, and a terminal `response.completed`, `response.incomplete`, or
`response.failed`. Sequence numbers order events; item IDs and output indexes remain
stable. The final response contains the full output and usage.

For Function Call, consume `response.function_call_arguments.delta` and
`response.function_call_arguments.done`. Gemini currently emits complete function
arguments in a single delta per call. Computer use actions follow the same format.
Code execution and media artifacts remain available in the output items'
`providerMetadata.google.parts`. Retain all final output items for the next turn.

Wallet settlement completes before the terminal response event is emitted. If
streaming or settlement fails, the endpoint emits `event: error` and closes without
a successful terminal event. Missing provider usage is retained for reconciliation.
Responses uses typed terminal events rather than Chat Completions' `[DONE]` marker.

`background: true` remains unsupported. This is a live HTTP stream subject to the
Vercel function time limit, not a persisted background task or a resumable stream.


## OpenAPI

````yaml openapi/google.yaml POST /responses
openapi: 3.1.0
info:
  title: Google models API
  version: 1.0.0
  description: GENGEN endpoints backed by Google Gemini models on Vertex AI.
servers:
  - url: https://gengen.farm/api/gengen/v1
    description: Production
security:
  - bearerAuth: []
paths:
  /responses:
    post:
      summary: Analyze multimodal input with Gemini
      description: >-
        Analyze multimodal input and use function calling, Code execution, or
        Computer use with synchronous or SSE output. See /tool-calling for
        replayable output items.
      operationId: createGoogleResponse
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/GoogleResponsesRequest'
            example:
              model: gemini-3.5-flash-lite
              input:
                - role: user
                  content:
                    - type: input_video
                      video_url: https://example.com/input.mp4
                      fps: 2
                    - type: input_text
                      text: Describe the sequence of actions in this video.
      responses:
        '200':
          description: An OpenAI-compatible Responses object or typed SSE stream.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ResponsesResult'
            text/event-stream:
              schema:
                type: string
                description: >-
                  Typed Responses SSE events. A terminal response event follows
                  successful settlement; transport or billing failures emit an
                  error event.
        default:
          $ref: '#/components/responses/Error'
components:
  schemas:
    GoogleResponsesRequest:
      type: object
      required:
        - model
        - input
      properties:
        model:
          type: string
          enum:
            - gemini-3.8-flash
            - gemini-3.7-flash
            - gemini-3.5-flash-lite
          description: Gemini model used for the response.
        input:
          oneOf:
            - type: string
            - type: array
              minItems: 1
              items:
                oneOf:
                  - $ref: '#/components/schemas/Message'
                  - $ref: '#/components/schemas/FunctionCallItem'
                  - $ref: '#/components/schemas/FunctionCallOutput'
          description: Text or multimodal conversation input.
        stream:
          type: boolean
          default: false
          description: >-
            Return typed Responses SSE events, including function argument
            deltas and terminal response usage.
        background:
          type: boolean
          const: false
          default: false
          description: Background execution is not supported.
        instructions:
          type: string
          description: High-level instructions sent as the Gemini system instruction.
        reasoning_effort:
          type: string
          enum:
            - minimal
            - low
            - medium
            - high
          description: >-
            Thinking level. `minimal` is supported by Gemini 3.5 Flash-Lite but
            not Gemini 3.8 Flash or Gemini 3.7 Flash. Gemini 3.8 Flash defaults
            to `medium` and rejects thinking budgets.
        temperature:
          type: number
          description: >-
            Sampling temperature. Custom values are ignored by Gemini 3.5
            Flash-Lite.
        top_p:
          type: number
          minimum: 0
          maximum: 1
          description: >-
            Nucleus-sampling probability threshold. Custom values are ignored by
            Gemini 3.5 Flash-Lite.
        top_k:
          type: number
          minimum: 1
          description: >-
            Limits sampling to the most likely K tokens. Custom values are
            ignored by Gemini 3.5 Flash-Lite.
        max_output_tokens:
          type: integer
          minimum: 1
          description: >-
            Maximum output tokens. `max_tokens` and `max_completion_tokens` are
            also accepted.
        max_tokens:
          type: integer
          minimum: 1
        stop:
          oneOf:
            - type: string
            - type: array
              items:
                type: string
          description: One or more sequences that stop generation.
        seed:
          type: integer
          description: Best-effort sampling seed.
        text:
          type: object
          description: Structured-output settings for supported Gemini models.
          properties:
            format:
              $ref: '#/components/schemas/ResponseFormat'
        tools:
          type: array
          description: >-
            Function declarations. Chat uses nested function objects; Responses
            uses flat function declarations. Built-in tools belong in
            providerOptions.google.tools.
          items:
            type: object
            additionalProperties: true
        tool_choice:
          description: >-
            auto, none, required, or a named function object. The function name
            must be declared in tools.
          oneOf:
            - type: string
              enum:
                - auto
                - none
                - required
            - type: object
              additionalProperties: true
        parallel_tool_calls:
          type: boolean
          const: true
          description: >-
            Omit or use true. Gemini can return multiple calls; false is
            rejected.
        providerOptions:
          $ref: '#/components/schemas/GoogleProviderOptions'
    ResponsesResult:
      type: object
      properties:
        id:
          type: string
        object:
          type: string
          const: response
        created_at:
          type: integer
        status:
          type: string
          enum:
            - completed
            - incomplete
            - failed
        model:
          type: string
        output:
          type: array
          items:
            type: object
            additionalProperties: true
        output_text:
          type: string
        usage:
          type: object
          additionalProperties: true
    Message:
      type: object
      required:
        - role
      properties:
        role:
          type: string
          enum:
            - system
            - developer
            - user
            - assistant
            - tool
          description: >-
            Message author. Assistant messages become Gemini `model` content.
            Gemini 3.5 Flash-Lite rejects a request whose final non-system
            message is `assistant`.
        content:
          oneOf:
            - type: 'null'
            - type: string
            - type: array
              items:
                $ref: '#/components/schemas/ContentPart'
        tool_calls:
          type: array
          items:
            type: object
            additionalProperties: true
          description: Assistant function calls. Preserve call IDs and extra_content.
        tool_call_id:
          type: string
          description: Matching call ID for a tool result.
        providerMetadata:
          type: object
          additionalProperties: true
          description: >-
            Preserve google.parts unchanged when replaying model output,
            including thought signatures and execution results.
    FunctionCallItem:
      type: object
      required:
        - type
        - call_id
        - name
        - arguments
      properties:
        type:
          type: string
          const: function_call
        call_id:
          type: string
        name:
          type: string
        arguments:
          type: string
        providerMetadata:
          type: object
          additionalProperties: true
    FunctionCallOutput:
      type: object
      required:
        - type
        - call_id
        - output
      properties:
        type:
          type: string
          const: function_call_output
        call_id:
          type: string
        output:
          description: >-
            JSON-string result, or content parts containing result text and
            screenshots.
          oneOf:
            - type: string
            - type: array
              items:
                $ref: '#/components/schemas/ContentPart'
    ResponseFormat:
      type: object
      required:
        - type
      properties:
        type:
          type: string
          enum:
            - text
            - json_object
            - json_schema
          description: >-
            Structured JSON modes are supported by both public Gemini text
            models.
        json_schema:
          type: object
          properties:
            name:
              type: string
            strict:
              type: boolean
            schema:
              type: object
              additionalProperties: true
          description: JSON Schema definition for `json_schema` output.
        schema:
          type: object
          additionalProperties: true
          description: Responses-style JSON Schema definition.
    GoogleProviderOptions:
      type: object
      properties:
        google:
          type: object
          properties:
            tools:
              type: array
              description: >-
                Native Google tools. Each entry defines functionDeclarations,
                codeExecution, or computerUse. Other built-in tools are
                rejected.
              items:
                type: object
                additionalProperties: true
              examples:
                - - codeExecution: {}
                - - computerUse:
                      environment: ENVIRONMENT_BROWSER
            toolConfig:
              type: object
              additionalProperties: true
              description: >-
                Google ToolConfig. Normalized tool_choice takes precedence.
                streamFunctionCallArguments=true is rejected; SSE emits complete
                call arguments.
            thinkingLevel:
              type: string
              enum:
                - MINIMAL
                - LOW
                - MEDIUM
                - HIGH
              description: >-
                Provider-native alias for `reasoning_effort`. `MINIMAL` is
                supported only by Gemini 3.5 Flash-Lite.
            generationConfig:
              type: object
              description: >-
                Advanced Vertex AI GenerationConfig fields. Normalized top-level
                fields take precedence.
              properties:
                mediaResolution:
                  type: string
                  enum:
                    - MEDIA_RESOLUTION_LOW
                    - MEDIA_RESOLUTION_MEDIUM
                    - MEDIA_RESOLUTION_HIGH
                responseLogprobs:
                  type: boolean
                logprobs:
                  type: integer
                  minimum: 1
                  maximum: 20
            safetySettings:
              type: array
              description: Vertex AI safety settings forwarded unchanged.
              items:
                type: object
                additionalProperties: true
    Error:
      type: object
      properties:
        error:
          type: object
          properties:
            code:
              type: string
            message:
              type: string
    ContentPart:
      oneOf:
        - $ref: '#/components/schemas/TextPart'
        - $ref: '#/components/schemas/ImagePart'
        - $ref: '#/components/schemas/VideoPart'
        - $ref: '#/components/schemas/AudioPart'
        - type: object
          required:
            - type
            - file_data
          properties:
            type:
              type: string
              const: input_file
            file_data:
              type: string
              description: >-
                Base64 data URL for an inline file, including CSV for Code
                execution.
    TextPart:
      type: object
      required:
        - type
        - text
      properties:
        type:
          type: string
          enum:
            - input_text
            - output_text
            - text
        text:
          type: string
          description: Text supplied to Gemini.
    ImagePart:
      type: object
      required:
        - type
        - image_url
      properties:
        type:
          type: string
          enum:
            - input_image
            - image_url
        image_url:
          oneOf:
            - type: string
            - type: object
              required:
                - url
              properties:
                url:
                  type: string
          description: >-
            Public or signed HTTPS URL, or Base64 data URL. Raw `gs://` URIs are
            not accepted.
    VideoPart:
      type: object
      required:
        - type
        - video_url
      properties:
        type:
          type: string
          const: input_video
        video_url:
          oneOf:
            - type: string
            - type: object
              required:
                - url
              properties:
                url:
                  type: string
          description: >-
            Public or signed HTTPS URL, or Base64 data URL. Raw `gs://` URIs are
            not accepted.
        fps:
          type: number
          exclusiveMinimum: 0
          maximum: 24
          default: 1
          description: Video sampling rate passed to Google.
        startOffset:
          oneOf:
            - type: number
            - type: string
          description: Clip start in seconds, such as `2` or `2s`.
        endOffset:
          oneOf:
            - type: number
            - type: string
          description: Clip end in seconds, such as `8.5` or `8.5s`.
        mediaResolution:
          type: string
          enum:
            - low
            - medium
            - high
          description: Per-video tokenization quality.
    AudioPart:
      type: object
      required:
        - type
        - audio_url
      properties:
        type:
          type: string
          const: input_audio
        audio_url:
          oneOf:
            - type: string
            - type: object
              required:
                - url
              properties:
                url:
                  type: string
          description: >-
            Public or signed HTTPS URL, or Base64 data URL. Raw `gs://` URIs are
            not accepted.
  responses:
    Error:
      description: The request could not be completed.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      bearerFormat: GENGEN API key
      description: A workspace API key beginning with `gengen_live_`.

````