POST /api/gengen/v1/chat/completions, including SSE streaming.
The /v1/chat/completions compatibility alias accepts the same requests.
Gemini gemini-3.8-flash, gemini-3.7-flash, and gemini-3.5-flash-lite also support
Google Code execution and Computer use. These tools are enabled per request.
Function calling
Send OpenAI-compatibletools and tool_choice. The model returns a structured
function call; your application executes the function and sends its result in the
next request. GENGEN does not execute user-defined functions.
This JavaScript example completes one round of function calls and sends all results
back. Production applications should continue the loop while calls remain and apply
their own execution permissions and maximum iteration count.
content: null and tool_calls. A normally
completed function-call turn uses finish_reason: "tool_calls". An interrupted
turn can have length or a provider failure reason; check it before executing a call.
For Gemini, tool_choice accepts auto, none, required, or
{"type":"function","function":{"name":"get_weather"}}.
Return a result for every call, matched by ID, including parallel calls with the same
function name. Gemini rejects parallel_tool_calls: false; handle all returned calls.
BytePlus forwards tool options and reasoning fields to the selected upstream model.
Gemini conversation metadata
Gemini assistant messages includeproviderMetadata.google.parts when needed. These
are the original ordered Gemini parts, including thought signatures, native call IDs,
code, execution results, and inline output bytes. Preserve them unchanged when
replaying the assistant message. They take precedence over reconstructed content.
Each Chat Completions tool call also carries extra_content.google.thought_signature
and extra_content.google.function_call_id when supplied by Google. Preserve these
if your client reconstructs messages from tool_calls instead of retaining the full
assistant message. Do not fabricate signatures or discard provider reasoning fields.
Streaming
Usestream: true on Chat Completions. Tool calls appear in
choices[0].delta.tool_calls with stable call IDs and zero-based indexes.
Gemini emits complete arguments for each call. Collect text and tool calls, and
concatenate delta.providerMetadata.google.parts arrays in stream order into the
assistant message’s providerMetadata.google.parts before replaying it.
Execution code, results, and graph bytes also arrive in these parts, even when the
chunk has no text. The stream ends with usage and data: [DONE] after successful
settlement. An error event without [DONE] is an incomplete stream.
Gemini’s native streamFunctionCallArguments: true option is rejected; the API does
not expose Google’s JSONPath argument fragments.
Gemini Code execution
Enable Google’s built-in Python execution tool:executableCode, codeExecutionResult, and any
inlineData graph output from the returned providerMetadata.google.parts.
Check the execution outcome; failed or timed-out code can still return a model
explanation. Code generation and execution results are preserved without requiring
text output.
For CSV or other supported inline files, use a content part such as
{"type":"input_file","file_data":"data:text/csv;base64,eCx5CjEsMg=="}.
Google Code execution requires inline file bytes rather than file URIs.
Gemini Computer use
Provide a screenshot and enable the built-in tool:tool_calls with the action name and arguments.
Your application supplies the browser or device, executes permitted actions, and
captures the new screenshot. GENGEN does not provision or control a computer.
Keep safety decisions in the call arguments. When Google requests confirmation,
ask the end user before executing the action. After actual confirmation, send the
provider’s safety_acknowledgement in the result JSON. GENGEN never supplies this
acknowledgement automatically.
Append the complete assistant message, then a tool result with the matching ID:
functionResponse.response and the screenshot as
functionResponse.parts. computerUse also accepts Google’s
excludedPredefinedFunctions and enablePromptInjectionDetection options.
ENVIRONMENT_MOBILE can be requested where supported by the selected Google model.
Responses API
Gemini also supports all three tool types through synchronous or streamingPOST /api/gengen/v1/responses. Use flat function declarations:
output items with type: "function_call", call_id, name, and
JSON-string arguments. Append all returned output items to your prior input,
then append one {"type":"function_call_output","call_id":"...","output":"..."}
for each call. Keep each item’s providerMetadata unchanged. output can also be
an array of text and screenshot content parts for Computer use.
Code and execution-result parts are carried in the corresponding assistant output
item’s providerMetadata.google.parts. output_text contains only the final visible
text. Set stream: true for typed Responses SSE events, including text and function
argument deltas. Read Streaming Responses
for the event lifecycle. background: true remains rejected. Send the full
conversation history; Gemini does not resolve
previous_response_id server-side.
BytePlus Responses remains available for the Seed models listed in its API reference;
Function Call on DeepSeek and GLM uses Chat Completions.
Usage and billing
Every model request is metered, including turns that return only a tool call and streamed completions. Gemini usage includes Code execution’s intermediate tool input tokens inprompt_tokens (Chat Completions) or input_tokens (Responses).
Cached prompt tokens retain their cache pricing; intermediate tool inputs use the
model’s ordinary input price. Code execution has no separate execution fee.
User-defined functions and browser execution happen in your application.
See Google’s Function calling,
Code execution, and
Computer use
documentation for provider-specific tool limitations.
Streaming model coverage
All 11 listed LLMs support streaming through Chat Completions. The six models with
a public Responses endpoint also support Responses SSE. BytePlus preserves its
native Responses events, including function argument fragments. Both providers
settle usage before delivering a terminal response event.