Skip to main content
Gemini 3.5 Flash-Lite is a cost-efficient multimodal model for high-volume, latency-sensitive text and media understanding workloads. Use minimal reasoning for extraction and other straightforward tasks. The model defaults to minimal; low, medium, and high are also supported. Custom temperature, top_p, and top_k values are ignored. frequency_penalty and presence_penalty are not supported, and the final non-system message must not use the assistant role.

Model ID

Pass this exact public ID in the top-level model field. Do not substitute an upstream provider alias.

Supported modes

  • chat_completions
  • responses_api
  • visual_understanding
  • text_generation

Inputs

  • Text
  • Image
  • Video
  • Audio

Outputs

  • Text

Capabilities

  • Not specified

Additional capabilities

  • Chat completions
  • Function calling
  • Code execution
  • Computer use
  • Responses API
  • Video understanding
  • SSE streaming
  • Structured output
  • Multimodal understanding
  • Configurable reasoning See Function calling and Gemini tools for tool declarations, result submission, and provider-specific capabilities.

Technical details

Availability and final customer prices can change. Check the Dashboard model catalog before production use.
See the API reference overview for authentication, request conventions, and response schemas.