Ultra-lightweight model with million-token context, optimized for speed and low latency, costing only $0.10 per million input tokens. It is suitable for edge computing and real-time interaction. The automatic caching mechanism offers a 75% cost reduction on cache hits.
Pricing
Input Modalities
- Text
- Vision
Output Modalities
- Text
Capabilities
- Tools
- Tool calling
- Structured outputs
- Long context
Try this model
Python
