Model card
Gemini 3.5 Flash Lite is engineered for developers building high-density, multi-agent architectures where latency and cost-per-token are the primary constraints. Unlike larger flagship models designed for broad reasoning, this model is optimized for 'agentic subtasks'—the granular, repetitive operations that occur within a larger workflow, such as data extraction, tool calling, or intent classification. It features a massive 1M token context window, allowing it to ingest large documentation sets or long conversation histories without losing coherence. For teams scaling agentic loops, Flash Lite offers a sweet spot: it provides the specialized instruction-following required for autonomous tool execution while maintaining the rapid inference speeds necessary to prevent bottlenecking complex pipelines. It is a strategic choice for developers moving from prototyping to production-scale agentic deployments.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page