gemini-3.5-flash-lite

by Google

Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing. The model supports text, image, video, audio, and PDF inputs, and is designed for high-volume agentic workflows, simple data extraction, and applications where latency and API cost are the primary constraints.

API Pricing

Input$0.3 / 1M tokens
Output$2.5 / 1M tokens

Specifications

Context window1,048,576 tokens
Modalitiestext, image, audio, video
Featuresthinking, tools, function_calling, structured_outputs, web, deepsearch, long_context
Endpointschat_completions, gemini_api, claude_api

Free version: gemini-3.5-flash-lite-free

More from Google

Use gemini-3.5-flash-lite via the AIHubMix unified API — one interface for every major LLM.