by Google
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing. The model supports text, image, video, audio, and PDF inputs, and is designed for high-volume agentic workflows, simple data extraction, and applications where latency and API cost are the primary constraints.
| Input | $0.3 / 1M tokens |
| Output | $2.5 / 1M tokens |
| Context window | 1,048,576 tokens |
| Modalities | text, image, audio, video |
| Features | thinking, tools, function_calling, structured_outputs, web, deepsearch, long_context |
| Endpoints | chat_completions, gemini_api, claude_api |
Free version: gemini-3.5-flash-lite-free
Use gemini-3.5-flash-lite via the AIHubMix unified API — one interface for every major LLM.