Skip to main content
POST
Create a Gemini chat completion
Use the OpenAI-compatible Chat Completions endpoint with gemini-3.8-flash, gemini-3.7-flash, or gemini-3.5-flash-lite.

Non-streaming request

The response uses the OpenAI-compatible choices and usage fields.

Streaming request

Set stream to true. GENGEN returns text/event-stream, emits partial Chat Completions chunks, includes final usage, and terminates the stream with data: [DONE].
This is unidirectional text streaming over SSE. It is separate from the bidirectional Gemini Live API.

Reasoning level

Gemini 3.8 Flash and Gemini 3.7 Flash accept low, medium, or high and default to medium. Gemini 3.5 Flash-Lite also accepts minimal, defaults to minimal, and recommends it for extraction and other straightforward, latency-sensitive tasks.

Gemini 3.5 Flash-Lite constraints

Gemini 3.5 Flash-Lite ignores custom temperature, top_p, and top_k values. It rejects frequency_penalty and presence_penalty, and the final non-system message must not use the assistant role.

Parameters

The API fixes candidateCount to 1 so the OpenAI-compatible response always contains one generated choice. Google may not produce identical output for repeated requests with the same seed. See the Gemini 3.8 Flash, Gemini 3.7 Flash and Gemini 3.5 Flash-Lite model pages for details. Gemini 3.8 Flash defaults to medium reasoning and rejects minimal, unsupported reasoning levels, and thinkingBudget overrides. Use low, medium, or high.

Function calling and built-in tools

Function calling, Code execution, and Computer use are supported for all three Gemini models. See Function calling and Gemini tools for complete requests, multi-turn result submission, thought signature preservation, streaming behavior, and billing.

Authorizations

Authorization
string
header
required

A workspace API key beginning with gengen_live_.

Body

application/json
model
enum<string>
required

Public Gemini text model ID.

Available options:
gemini-3.8-flash,
gemini-3.7-flash,
gemini-3.5-flash-lite
messages
object[]
required

Conversation messages. System and developer messages become the Gemini system instruction.

Minimum array length: 1
reasoning_effort
enum<string>

Gemini thinking level. minimal is supported by Gemini 3.5 Flash-Lite but not Gemini 3.8 Flash or Gemini 3.7 Flash. Gemini 3.8 Flash defaults to medium and rejects thinking budgets.

Available options:
minimal,
low,
medium,
high
temperature
number

Sampling temperature. Higher values increase variation. Custom values are ignored by Gemini 3.5 Flash-Lite.

top_p
number

Nucleus-sampling probability threshold. Custom values are ignored by Gemini 3.5 Flash-Lite.

Required range: 0 <= x <= 1
top_k
number

Limits sampling to the most likely K tokens. Custom values are ignored by Gemini 3.5 Flash-Lite.

Required range: x >= 1
max_tokens
integer

Maximum output token count. max_completion_tokens is an accepted alias.

Required range: x >= 1
max_completion_tokens
integer

Maximum output token count.

Required range: x >= 1
stop

One or more sequences that stop generation.

frequency_penalty
number

Penalizes tokens according to how often they already appear. Gemini 3.5 Flash-Lite rejects this parameter.

presence_penalty
number

Penalizes tokens that have already appeared at least once. Gemini 3.5 Flash-Lite rejects this parameter.

seed
integer

Best-effort sampling seed. Identical outputs are not guaranteed.

response_format
object
stream
boolean
default:false

Return Server-Sent Events when true.

stream_options
object

Streaming compatibility options.

tools
object[]

Function declarations. Chat uses nested function objects; Responses uses flat function declarations. Built-in tools belong in providerOptions.google.tools.

tool_choice

auto, none, required, or a named function object. The function name must be declared in tools.

Available options:
auto,
none,
required
parallel_tool_calls
boolean

Omit or use true. Gemini can return multiple calls; false is rejected.

providerOptions
object

Response

An OpenAI-compatible chat completion or SSE stream.

id
string
object
string
Allowed value: "chat.completion"
created
integer
model
string
choices
object[]
usage
object