Model card
Gemini 3.5 Flash Lite (Batch) is engineered for developers building high-throughput, agentic workflows where latency and cost-efficiency are the primary constraints. Unlike larger flagship models designed for broad reasoning, this model is optimized as a specialized subagent. It excels at executing granular, discrete tasks—such as data extraction, classification, or tool-calling—within a larger multi-agent orchestration. By utilizing the batch processing mode, developers can significantly reduce costs for non-real-time workloads while maintaining a massive 1M token context window. This makes it an ideal choice for processing large datasets or managing background asynchronous tasks in complex pipelines. If your architecture requires a 'worker' model to handle high-volume, repetitive logic without the overhead of a heavy-duty LLM, this version provides the necessary speed and scale for production-grade automation.
Model files and versions
Download this model
How to use
- 01Step 1
Read the model card and source information.
- 02Step 2
Start with a small, non-sensitive evaluation.
- 03Step 3
Review quality, licensing and usage limits.
- 04Step 4
Adopt it only after validation.
Discussions
Use this space to keep checking source information, usage experience and maintenance status.
Open source page