> ## Documentation Index
> Fetch the complete documentation index at: https://docs.gengen.farm/llms.txt
> Use this file to discover all available pages before exploring further.

# Chat completions with Gemini

> Generate synchronous or streaming text with Gemini 3.8 Flash, Gemini 3.7 Flash, or Gemini 3.5 Flash-Lite through the GENGEN Chat Completions API.

Use the OpenAI-compatible Chat Completions endpoint with `gemini-3.8-flash`, `gemini-3.7-flash`, or `gemini-3.5-flash-lite`.

| Property | Value |
| - | - |
| Method | `POST` |
| Endpoint | `/api/gengen/v1/chat/completions` |
| Models | `gemini-3.8-flash`, `gemini-3.7-flash`, `gemini-3.5-flash-lite` |
| Streaming | Server-Sent Events (SSE) |

## Non-streaming request

```bash theme={"dark"}
curl --request POST \
  --url https://gengen.farm/api/gengen/v1/chat/completions \
  --header 'Authorization: Bearer gengen_live_xxxxxxxxxxxxxxxx' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "gemini-3.5-flash-lite",
    "messages": [
      {
        "role": "system",
        "content": "You are a concise technical assistant."
      },
      {
        "role": "user",
        "content": "Explain idempotency in two sentences."
      }
    ],
    "reasoning_effort": "minimal",
    "max_tokens": 512,
    "stream": false
  }'
```

The response uses the OpenAI-compatible `choices` and `usage` fields.

```json theme={"dark"}
{
  "id": "google-response-id",
  "object": "chat.completion",
  "model": "gemini-3.5-flash-lite",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Idempotency means repeating the same operation produces the same intended effect. It prevents retries from creating duplicate changes."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 32,
    "completion_tokens": 28,
    "total_tokens": 60
  }
}
```

## Streaming request

Set `stream` to `true`. GENGEN returns `text/event-stream`, emits partial Chat Completions chunks, includes final usage, and terminates the stream with `data: [DONE]`.

```bash theme={"dark"}
curl --no-buffer --request POST \
  --url https://gengen.farm/api/gengen/v1/chat/completions \
  --header 'Authorization: Bearer gengen_live_xxxxxxxxxxxxxxxx' \
  --header 'Content-Type: application/json' \
  --data '{
    "model": "gemini-3.8-flash",
    "messages": [
      {
        "role": "user",
        "content": "Write a one-line product description."
      }
    ],
    "stream": true,
    "stream_options": {
      "include_usage": true
    }
  }'
```

```text theme={"dark"}
data: {"id":"google-response-id","object":"chat.completion.chunk","model":"gemini-3.8-flash","choices":[{"index":0,"delta":{"content":"A fast"},"finish_reason":null}]}

data: {"id":"google-response-id","object":"chat.completion.chunk","model":"gemini-3.8-flash","choices":[],"usage":{"prompt_tokens":16,"completion_tokens":8,"total_tokens":24}}

data: [DONE]
```

<Note>
  This is unidirectional text streaming over SSE. It is separate from the bidirectional Gemini Live API.
</Note>

## Reasoning level

Gemini 3.8 Flash and Gemini 3.7 Flash accept `low`, `medium`, or `high` and default to `medium`. Gemini 3.5 Flash-Lite also accepts `minimal`, defaults to `minimal`, and recommends it for extraction and other straightforward, latency-sensitive tasks.

## Gemini 3.5 Flash-Lite constraints

Gemini 3.5 Flash-Lite ignores custom `temperature`, `top_p`, and `top_k` values. It rejects `frequency_penalty` and `presence_penalty`, and the final non-system message must not use the `assistant` role.

## Parameters

| Parameter | Type | Description |
| - | - | - |
| `messages` | array | Required conversation history. `system` and `developer` messages become the Gemini system instruction; text, image, video, and audio parts are supported. |
| `reasoning_effort` | string | `low`, `medium`, or `high` for Gemini 3.8 Flash and Gemini 3.7 Flash; Gemini 3.5 Flash-Lite also supports `minimal`. |
| `temperature` | number | Sampling temperature. Gemini 3.5 Flash-Lite ignores custom values. |
| `top_p` | number | Nucleus-sampling threshold from `0` to `1`. Gemini 3.5 Flash-Lite ignores custom values. |
| `top_k` | number | Restricts sampling to the most likely K tokens. Gemini 3.5 Flash-Lite ignores custom values. |
| `max_tokens` / `max_completion_tokens` | integer | Maximum generated tokens. |
| `stop` | string or string\[] | One or more stop sequences. |
| `frequency_penalty` | number | Penalizes repeated token frequency. Not supported by Gemini 3.5 Flash-Lite. |
| `presence_penalty` | number | Penalizes tokens already present in the output. Not supported by Gemini 3.5 Flash-Lite. |
| `seed` | integer | Best-effort deterministic sampling seed. |
| `response_format` | object | Use `json_object` or `json_schema` for structured output. |
| `stream` | boolean | Enables SSE streaming. |
| `providerOptions.google` | object | Advanced Google `generationConfig`, `safetySettings`, `tools`, and `toolConfig`. Normalized parameters above take precedence. |

The API fixes `candidateCount` to `1` so the OpenAI-compatible response always contains one generated choice. Google may not produce identical output for repeated requests with the same `seed`.

See the [Gemini 3.8 Flash](/models/gemini-3.8-flash), [Gemini 3.7 Flash](/models/gemini-3.7-flash) and [Gemini 3.5 Flash-Lite](/models/gemini-3.5-flash-lite) model pages for details.

Gemini 3.8 Flash defaults to `medium` reasoning and rejects `minimal`, unsupported reasoning levels, and `thinkingBudget` overrides. Use `low`, `medium`, or `high`.

## Function calling and built-in tools

Function calling, Code execution, and Computer use are supported for all three Gemini models.
See [Function calling and Gemini tools](/tool-calling) for complete requests, multi-turn
result submission, thought signature preservation, streaming behavior, and billing.


## OpenAPI

````yaml openapi/google.yaml POST /chat/completions
openapi: 3.1.0
info:
  title: Google models API
  version: 1.0.0
  description: GENGEN endpoints backed by Google Gemini models on Vertex AI.
servers:
  - url: https://gengen.farm/api/gengen/v1
    description: Production
security:
  - bearerAuth: []
paths:
  /chat/completions:
    post:
      summary: Create a Gemini chat completion
      description: >-
        Generate text, function calls, code execution results, or computer
        actions with a supported Gemini model. See /tool-calling for multi-turn
        examples.
      operationId: createGoogleChatCompletion
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/GoogleChatRequest'
            example:
              model: gemini-3.5-flash-lite
              messages:
                - role: system
                  content: You are a concise technical assistant.
                - role: user
                  content: Explain idempotency in two sentences.
              reasoning_effort: minimal
              max_tokens: 512
              stream: false
      responses:
        '200':
          description: An OpenAI-compatible chat completion or SSE stream.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ChatCompletion'
        default:
          $ref: '#/components/responses/Error'
components:
  schemas:
    GoogleChatRequest:
      type: object
      required:
        - model
        - messages
      properties:
        model:
          type: string
          enum:
            - gemini-3.8-flash
            - gemini-3.7-flash
            - gemini-3.5-flash-lite
          description: Public Gemini text model ID.
        messages:
          type: array
          minItems: 1
          description: >-
            Conversation messages. System and developer messages become the
            Gemini system instruction.
          items:
            $ref: '#/components/schemas/Message'
        reasoning_effort:
          type: string
          enum:
            - minimal
            - low
            - medium
            - high
          description: >-
            Gemini thinking level. `minimal` is supported by Gemini 3.5
            Flash-Lite but not Gemini 3.8 Flash or Gemini 3.7 Flash. Gemini 3.8
            Flash defaults to `medium` and rejects thinking budgets.
        temperature:
          type: number
          description: >-
            Sampling temperature. Higher values increase variation. Custom
            values are ignored by Gemini 3.5 Flash-Lite.
        top_p:
          type: number
          minimum: 0
          maximum: 1
          description: >-
            Nucleus-sampling probability threshold. Custom values are ignored by
            Gemini 3.5 Flash-Lite.
        top_k:
          type: number
          minimum: 1
          description: >-
            Limits sampling to the most likely K tokens. Custom values are
            ignored by Gemini 3.5 Flash-Lite.
        max_tokens:
          type: integer
          minimum: 1
          description: >-
            Maximum output token count. `max_completion_tokens` is an accepted
            alias.
        max_completion_tokens:
          type: integer
          minimum: 1
          description: Maximum output token count.
        stop:
          oneOf:
            - type: string
            - type: array
              items:
                type: string
          description: One or more sequences that stop generation.
        frequency_penalty:
          type: number
          description: >-
            Penalizes tokens according to how often they already appear. Gemini
            3.5 Flash-Lite rejects this parameter.
        presence_penalty:
          type: number
          description: >-
            Penalizes tokens that have already appeared at least once. Gemini
            3.5 Flash-Lite rejects this parameter.
        seed:
          type: integer
          description: Best-effort sampling seed. Identical outputs are not guaranteed.
        response_format:
          $ref: '#/components/schemas/ResponseFormat'
        stream:
          type: boolean
          default: false
          description: Return Server-Sent Events when true.
        stream_options:
          type: object
          description: Streaming compatibility options.
          properties:
            include_usage:
              type: boolean
              description: Include usage in the final stream event.
        tools:
          type: array
          description: >-
            Function declarations. Chat uses nested function objects; Responses
            uses flat function declarations. Built-in tools belong in
            providerOptions.google.tools.
          items:
            type: object
            additionalProperties: true
        tool_choice:
          description: >-
            auto, none, required, or a named function object. The function name
            must be declared in tools.
          oneOf:
            - type: string
              enum:
                - auto
                - none
                - required
            - type: object
              additionalProperties: true
        parallel_tool_calls:
          type: boolean
          const: true
          description: >-
            Omit or use true. Gemini can return multiple calls; false is
            rejected.
        providerOptions:
          $ref: '#/components/schemas/GoogleProviderOptions'
    ChatCompletion:
      type: object
      properties:
        id:
          type: string
        object:
          type: string
          const: chat.completion
        created:
          type: integer
        model:
          type: string
        choices:
          type: array
          items:
            type: object
            additionalProperties: true
        usage:
          type: object
          additionalProperties: true
    Message:
      type: object
      required:
        - role
      properties:
        role:
          type: string
          enum:
            - system
            - developer
            - user
            - assistant
            - tool
          description: >-
            Message author. Assistant messages become Gemini `model` content.
            Gemini 3.5 Flash-Lite rejects a request whose final non-system
            message is `assistant`.
        content:
          oneOf:
            - type: 'null'
            - type: string
            - type: array
              items:
                $ref: '#/components/schemas/ContentPart'
        tool_calls:
          type: array
          items:
            type: object
            additionalProperties: true
          description: Assistant function calls. Preserve call IDs and extra_content.
        tool_call_id:
          type: string
          description: Matching call ID for a tool result.
        providerMetadata:
          type: object
          additionalProperties: true
          description: >-
            Preserve google.parts unchanged when replaying model output,
            including thought signatures and execution results.
    ResponseFormat:
      type: object
      required:
        - type
      properties:
        type:
          type: string
          enum:
            - text
            - json_object
            - json_schema
          description: >-
            Structured JSON modes are supported by both public Gemini text
            models.
        json_schema:
          type: object
          properties:
            name:
              type: string
            strict:
              type: boolean
            schema:
              type: object
              additionalProperties: true
          description: JSON Schema definition for `json_schema` output.
        schema:
          type: object
          additionalProperties: true
          description: Responses-style JSON Schema definition.
    GoogleProviderOptions:
      type: object
      properties:
        google:
          type: object
          properties:
            tools:
              type: array
              description: >-
                Native Google tools. Each entry defines functionDeclarations,
                codeExecution, or computerUse. Other built-in tools are
                rejected.
              items:
                type: object
                additionalProperties: true
              examples:
                - - codeExecution: {}
                - - computerUse:
                      environment: ENVIRONMENT_BROWSER
            toolConfig:
              type: object
              additionalProperties: true
              description: >-
                Google ToolConfig. Normalized tool_choice takes precedence.
                streamFunctionCallArguments=true is rejected; SSE emits complete
                call arguments.
            thinkingLevel:
              type: string
              enum:
                - MINIMAL
                - LOW
                - MEDIUM
                - HIGH
              description: >-
                Provider-native alias for `reasoning_effort`. `MINIMAL` is
                supported only by Gemini 3.5 Flash-Lite.
            generationConfig:
              type: object
              description: >-
                Advanced Vertex AI GenerationConfig fields. Normalized top-level
                fields take precedence.
              properties:
                mediaResolution:
                  type: string
                  enum:
                    - MEDIA_RESOLUTION_LOW
                    - MEDIA_RESOLUTION_MEDIUM
                    - MEDIA_RESOLUTION_HIGH
                responseLogprobs:
                  type: boolean
                logprobs:
                  type: integer
                  minimum: 1
                  maximum: 20
            safetySettings:
              type: array
              description: Vertex AI safety settings forwarded unchanged.
              items:
                type: object
                additionalProperties: true
    Error:
      type: object
      properties:
        error:
          type: object
          properties:
            code:
              type: string
            message:
              type: string
    ContentPart:
      oneOf:
        - $ref: '#/components/schemas/TextPart'
        - $ref: '#/components/schemas/ImagePart'
        - $ref: '#/components/schemas/VideoPart'
        - $ref: '#/components/schemas/AudioPart'
        - type: object
          required:
            - type
            - file_data
          properties:
            type:
              type: string
              const: input_file
            file_data:
              type: string
              description: >-
                Base64 data URL for an inline file, including CSV for Code
                execution.
    TextPart:
      type: object
      required:
        - type
        - text
      properties:
        type:
          type: string
          enum:
            - input_text
            - output_text
            - text
        text:
          type: string
          description: Text supplied to Gemini.
    ImagePart:
      type: object
      required:
        - type
        - image_url
      properties:
        type:
          type: string
          enum:
            - input_image
            - image_url
        image_url:
          oneOf:
            - type: string
            - type: object
              required:
                - url
              properties:
                url:
                  type: string
          description: >-
            Public or signed HTTPS URL, or Base64 data URL. Raw `gs://` URIs are
            not accepted.
    VideoPart:
      type: object
      required:
        - type
        - video_url
      properties:
        type:
          type: string
          const: input_video
        video_url:
          oneOf:
            - type: string
            - type: object
              required:
                - url
              properties:
                url:
                  type: string
          description: >-
            Public or signed HTTPS URL, or Base64 data URL. Raw `gs://` URIs are
            not accepted.
        fps:
          type: number
          exclusiveMinimum: 0
          maximum: 24
          default: 1
          description: Video sampling rate passed to Google.
        startOffset:
          oneOf:
            - type: number
            - type: string
          description: Clip start in seconds, such as `2` or `2s`.
        endOffset:
          oneOf:
            - type: number
            - type: string
          description: Clip end in seconds, such as `8.5` or `8.5s`.
        mediaResolution:
          type: string
          enum:
            - low
            - medium
            - high
          description: Per-video tokenization quality.
    AudioPart:
      type: object
      required:
        - type
        - audio_url
      properties:
        type:
          type: string
          const: input_audio
        audio_url:
          oneOf:
            - type: string
            - type: object
              required:
                - url
              properties:
                url:
                  type: string
          description: >-
            Public or signed HTTPS URL, or Base64 data URL. Raw `gs://` URIs are
            not accepted.
  responses:
    Error:
      description: The request could not be completed.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      bearerFormat: GENGEN API key
      description: A workspace API key beginning with `gengen_live_`.

````