Skip to main content
All currently listed Seed, DeepSeek, GLM, and Gemini LLMs support function calling through POST /api/gengen/v1/chat/completions, including SSE streaming. The /v1/chat/completions compatibility alias accepts the same requests. Gemini gemini-3.8-flash, gemini-3.7-flash, and gemini-3.5-flash-lite also support Google Code execution and Computer use. These tools are enabled per request.

Function calling

Send OpenAI-compatible tools and tool_choice. The model returns a structured function call; your application executes the function and sends its result in the next request. GENGEN does not execute user-defined functions. This JavaScript example completes one round of function calls and sends all results back. Production applications should continue the loop while calls remain and apply their own execution permissions and maximum iteration count.
A function-only assistant message has content: null and tool_calls. A normally completed function-call turn uses finish_reason: "tool_calls". An interrupted turn can have length or a provider failure reason; check it before executing a call. For Gemini, tool_choice accepts auto, none, required, or {"type":"function","function":{"name":"get_weather"}}. Return a result for every call, matched by ID, including parallel calls with the same function name. Gemini rejects parallel_tool_calls: false; handle all returned calls. BytePlus forwards tool options and reasoning fields to the selected upstream model.

Gemini conversation metadata

Gemini assistant messages include providerMetadata.google.parts when needed. These are the original ordered Gemini parts, including thought signatures, native call IDs, code, execution results, and inline output bytes. Preserve them unchanged when replaying the assistant message. They take precedence over reconstructed content. Each Chat Completions tool call also carries extra_content.google.thought_signature and extra_content.google.function_call_id when supplied by Google. Preserve these if your client reconstructs messages from tool_calls instead of retaining the full assistant message. Do not fabricate signatures or discard provider reasoning fields.

Streaming

Use stream: true on Chat Completions. Tool calls appear in choices[0].delta.tool_calls with stable call IDs and zero-based indexes. Gemini emits complete arguments for each call. Collect text and tool calls, and concatenate delta.providerMetadata.google.parts arrays in stream order into the assistant message’s providerMetadata.google.parts before replaying it. Execution code, results, and graph bytes also arrive in these parts, even when the chunk has no text. The stream ends with usage and data: [DONE] after successful settlement. An error event without [DONE] is an incomplete stream. Gemini’s native streamFunctionCallArguments: true option is rejected; the API does not expose Google’s JSONPath argument fragments.

Gemini Code execution

Enable Google’s built-in Python execution tool:
Google runs the code. Read executableCode, codeExecutionResult, and any inlineData graph output from the returned providerMetadata.google.parts. Check the execution outcome; failed or timed-out code can still return a model explanation. Code generation and execution results are preserved without requiring text output. For CSV or other supported inline files, use a content part such as {"type":"input_file","file_data":"data:text/csv;base64,eCx5CjEsMg=="}. Google Code execution requires inline file bytes rather than file URIs.

Gemini Computer use

Provide a screenshot and enable the built-in tool:
Computer actions return as tool_calls with the action name and arguments. Your application supplies the browser or device, executes permitted actions, and captures the new screenshot. GENGEN does not provision or control a computer. Keep safety decisions in the call arguments. When Google requests confirmation, ask the end user before executing the action. After actual confirmation, send the provider’s safety_acknowledgement in the result JSON. GENGEN never supplies this acknowledgement automatically. Append the complete assistant message, then a tool result with the matching ID:
Gemini receives the JSON as functionResponse.response and the screenshot as functionResponse.parts. computerUse also accepts Google’s excludedPredefinedFunctions and enablePromptInjectionDetection options. ENVIRONMENT_MOBILE can be requested where supported by the selected Google model.

Responses API

Gemini also supports all three tool types through synchronous or streaming POST /api/gengen/v1/responses. Use flat function declarations:
Calls appear as output items with type: "function_call", call_id, name, and JSON-string arguments. Append all returned output items to your prior input, then append one {"type":"function_call_output","call_id":"...","output":"..."} for each call. Keep each item’s providerMetadata unchanged. output can also be an array of text and screenshot content parts for Computer use. Code and execution-result parts are carried in the corresponding assistant output item’s providerMetadata.google.parts. output_text contains only the final visible text. Set stream: true for typed Responses SSE events, including text and function argument deltas. Read Streaming Responses for the event lifecycle. background: true remains rejected. Send the full conversation history; Gemini does not resolve previous_response_id server-side. BytePlus Responses remains available for the Seed models listed in its API reference; Function Call on DeepSeek and GLM uses Chat Completions.

Usage and billing

Every model request is metered, including turns that return only a tool call and streamed completions. Gemini usage includes Code execution’s intermediate tool input tokens in prompt_tokens (Chat Completions) or input_tokens (Responses). Cached prompt tokens retain their cache pricing; intermediate tool inputs use the model’s ordinary input price. Code execution has no separate execution fee. User-defined functions and browser execution happen in your application. See Google’s Function calling, Code execution, and Computer use documentation for provider-specific tool limitations.

Streaming model coverage

All 11 listed LLMs support streaming through Chat Completions. The six models with a public Responses endpoint also support Responses SSE. BytePlus preserves its native Responses events, including function argument fragments. Both providers settle usage before delivering a terminal response event.