curl --request POST \
--url https://gengen.farm/api/gengen/v1/chat/completions \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "gemini-3.5-flash-lite",
"messages": [
{
"role": "system",
"content": "You are a concise technical assistant."
},
{
"role": "user",
"content": "Explain idempotency in two sentences."
}
]
}
'import requests
url = "https://gengen.farm/api/gengen/v1/chat/completions"
payload = {
"model": "gemini-3.5-flash-lite",
"messages": [
{
"role": "system",
"content": "You are a concise technical assistant."
},
{
"role": "user",
"content": "Explain idempotency in two sentences."
}
]
}
headers = {
"Authorization": "Bearer <token>",
"Content-Type": "application/json"
}
response = requests.post(url, json=payload, headers=headers)
print(response.text)const options = {
method: 'POST',
headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},
body: JSON.stringify({
model: 'gemini-3.5-flash-lite',
messages: [
{role: 'system', content: 'You are a concise technical assistant.'},
{role: 'user', content: 'Explain idempotency in two sentences.'}
]
})
};
fetch('https://gengen.farm/api/gengen/v1/chat/completions', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"id": "<string>",
"object": "chat.completion",
"created": 123,
"model": "<string>",
"choices": [
{}
],
"usage": {}
}{
"error": {
"code": "<string>",
"message": "<string>"
}
}Chat completions with Gemini
Generate synchronous or streaming text with Gemini 3.8 Flash, Gemini 3.7 Flash, or Gemini 3.5 Flash-Lite through the GENGEN Chat Completions API.
curl --request POST \
--url https://gengen.farm/api/gengen/v1/chat/completions \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"model": "gemini-3.5-flash-lite",
"messages": [
{
"role": "system",
"content": "You are a concise technical assistant."
},
{
"role": "user",
"content": "Explain idempotency in two sentences."
}
]
}
'import requests
url = "https://gengen.farm/api/gengen/v1/chat/completions"
payload = {
"model": "gemini-3.5-flash-lite",
"messages": [
{
"role": "system",
"content": "You are a concise technical assistant."
},
{
"role": "user",
"content": "Explain idempotency in two sentences."
}
]
}
headers = {
"Authorization": "Bearer <token>",
"Content-Type": "application/json"
}
response = requests.post(url, json=payload, headers=headers)
print(response.text)const options = {
method: 'POST',
headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},
body: JSON.stringify({
model: 'gemini-3.5-flash-lite',
messages: [
{role: 'system', content: 'You are a concise technical assistant.'},
{role: 'user', content: 'Explain idempotency in two sentences.'}
]
})
};
fetch('https://gengen.farm/api/gengen/v1/chat/completions', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"id": "<string>",
"object": "chat.completion",
"created": 123,
"model": "<string>",
"choices": [
{}
],
"usage": {}
}{
"error": {
"code": "<string>",
"message": "<string>"
}
}gemini-3.8-flash, gemini-3.7-flash, or gemini-3.5-flash-lite.
| Property | Value |
|---|---|
| Method | POST |
| Endpoint | /api/gengen/v1/chat/completions |
| Models | gemini-3.8-flash, gemini-3.7-flash, gemini-3.5-flash-lite |
| Streaming | Server-Sent Events (SSE) |
Non-streaming request
curl --request POST \
--url https://gengen.farm/api/gengen/v1/chat/completions \
--header 'Authorization: Bearer gengen_live_xxxxxxxxxxxxxxxx' \
--header 'Content-Type: application/json' \
--data '{
"model": "gemini-3.5-flash-lite",
"messages": [
{
"role": "system",
"content": "You are a concise technical assistant."
},
{
"role": "user",
"content": "Explain idempotency in two sentences."
}
],
"reasoning_effort": "minimal",
"max_tokens": 512,
"stream": false
}'
choices and usage fields.
{
"id": "google-response-id",
"object": "chat.completion",
"model": "gemini-3.5-flash-lite",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Idempotency means repeating the same operation produces the same intended effect. It prevents retries from creating duplicate changes."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 32,
"completion_tokens": 28,
"total_tokens": 60
}
}
Streaming request
Setstream to true. GENGEN returns text/event-stream, emits partial Chat Completions chunks, includes final usage, and terminates the stream with data: [DONE].
curl --no-buffer --request POST \
--url https://gengen.farm/api/gengen/v1/chat/completions \
--header 'Authorization: Bearer gengen_live_xxxxxxxxxxxxxxxx' \
--header 'Content-Type: application/json' \
--data '{
"model": "gemini-3.8-flash",
"messages": [
{
"role": "user",
"content": "Write a one-line product description."
}
],
"stream": true,
"stream_options": {
"include_usage": true
}
}'
data: {"id":"google-response-id","object":"chat.completion.chunk","model":"gemini-3.8-flash","choices":[{"index":0,"delta":{"content":"A fast"},"finish_reason":null}]}
data: {"id":"google-response-id","object":"chat.completion.chunk","model":"gemini-3.8-flash","choices":[],"usage":{"prompt_tokens":16,"completion_tokens":8,"total_tokens":24}}
data: [DONE]
Reasoning level
Gemini 3.8 Flash and Gemini 3.7 Flash acceptlow, medium, or high and default to medium. Gemini 3.5 Flash-Lite also accepts minimal, defaults to minimal, and recommends it for extraction and other straightforward, latency-sensitive tasks.
Gemini 3.5 Flash-Lite constraints
Gemini 3.5 Flash-Lite ignores customtemperature, top_p, and top_k values. It rejects frequency_penalty and presence_penalty, and the final non-system message must not use the assistant role.
Parameters
| Parameter | Type | Description |
|---|---|---|
messages | array | Required conversation history. system and developer messages become the Gemini system instruction; text, image, video, and audio parts are supported. |
reasoning_effort | string | low, medium, or high for Gemini 3.8 Flash and Gemini 3.7 Flash; Gemini 3.5 Flash-Lite also supports minimal. |
temperature | number | Sampling temperature. Gemini 3.5 Flash-Lite ignores custom values. |
top_p | number | Nucleus-sampling threshold from 0 to 1. Gemini 3.5 Flash-Lite ignores custom values. |
top_k | number | Restricts sampling to the most likely K tokens. Gemini 3.5 Flash-Lite ignores custom values. |
max_tokens / max_completion_tokens | integer | Maximum generated tokens. |
stop | string or string[] | One or more stop sequences. |
frequency_penalty | number | Penalizes repeated token frequency. Not supported by Gemini 3.5 Flash-Lite. |
presence_penalty | number | Penalizes tokens already present in the output. Not supported by Gemini 3.5 Flash-Lite. |
seed | integer | Best-effort deterministic sampling seed. |
response_format | object | Use json_object or json_schema for structured output. |
stream | boolean | Enables SSE streaming. |
providerOptions.google | object | Advanced Google generationConfig, safetySettings, tools, and toolConfig. Normalized parameters above take precedence. |
candidateCount to 1 so the OpenAI-compatible response always contains one generated choice. Google may not produce identical output for repeated requests with the same seed.
See the Gemini 3.8 Flash, Gemini 3.7 Flash and Gemini 3.5 Flash-Lite model pages for details.
Gemini 3.8 Flash defaults to medium reasoning and rejects minimal, unsupported reasoning levels, and thinkingBudget overrides. Use low, medium, or high.
Function calling and built-in tools
Function calling, Code execution, and Computer use are supported for all three Gemini models. See Function calling and Gemini tools for complete requests, multi-turn result submission, thought signature preservation, streaming behavior, and billing.Authorizations
A workspace API key beginning with gengen_live_.
Body
Public Gemini text model ID.
gemini-3.8-flash, gemini-3.7-flash, gemini-3.5-flash-lite Conversation messages. System and developer messages become the Gemini system instruction.
1Show child attributes
Show child attributes
Gemini thinking level. minimal is supported by Gemini 3.5 Flash-Lite but not Gemini 3.8 Flash or Gemini 3.7 Flash. Gemini 3.8 Flash defaults to medium and rejects thinking budgets.
minimal, low, medium, high Sampling temperature. Higher values increase variation. Custom values are ignored by Gemini 3.5 Flash-Lite.
Nucleus-sampling probability threshold. Custom values are ignored by Gemini 3.5 Flash-Lite.
0 <= x <= 1Limits sampling to the most likely K tokens. Custom values are ignored by Gemini 3.5 Flash-Lite.
x >= 1Maximum output token count. max_completion_tokens is an accepted alias.
x >= 1Maximum output token count.
x >= 1One or more sequences that stop generation.
Penalizes tokens according to how often they already appear. Gemini 3.5 Flash-Lite rejects this parameter.
Penalizes tokens that have already appeared at least once. Gemini 3.5 Flash-Lite rejects this parameter.
Best-effort sampling seed. Identical outputs are not guaranteed.
Show child attributes
Show child attributes
Return Server-Sent Events when true.
Streaming compatibility options.
Show child attributes
Show child attributes
Function declarations. Chat uses nested function objects; Responses uses flat function declarations. Built-in tools belong in providerOptions.google.tools.
auto, none, required, or a named function object. The function name must be declared in tools.
auto, none, required Omit or use true. Gemini can return multiple calls; false is rejected.
Show child attributes
Show child attributes