Sol API pricing changes increase the cost of generating output tokens
The updated Sol API rate cards necessitate a redesign of LLM agent workflows to avoid budget overruns. Developers executing long-form generation or complex reasoning chains should analyze these costs before expanding their deployments.
A sharp divide now exists between input and output pricing. Processing prompts costs $4.00 per million tokens, whereas generating responses costs $20.00 per million tokens. This means output is five times more expensive than input. While RAG users typically see higher bills from large context windows in the input, those using iterative "Chain of Thought" reasoning will find their budgets depleted by the high cost of model output.
Consider a scenario with a 10,000 token prompt and a 1,000 token response. The input costs $0.04 and the output costs $0.02. These figures escalate quickly in recursive loops where agents re-process entire histories after tool calls.
This model may act as a middle ground for those currently using more expensive frontier models. However, users of fine-tuned models should use aggressive prompt engineering to limit output. Wordy system instructions often cause the model to narrate actions, which wastes money at a $20 per million rate.
Production workflows require immediate implementation of token usage trackers to prevent expensive errors from repetitive patterns or runaway loops. One strategy to reduce costs is using cheaper models for data extraction or summarization and reserving Sol for the final high-reasoning phase.
{
"model": "sol-api-latest",
"usage_estimate": {
"input_rate_per_1m": 4.00,
"output_rate_per_1m": 20.00,
"currency": "USD"
}
}All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Those costs are brutal. Are you caching responses or just looking for a cheaper model? If you're running heavy reasoning chains or long-form generation tasks, you'll want to scrutinize these figures before scaling up your deployment.
Ridiculous pricing. I started batching requests to avoid those micro-transaction fees. You can also reduce output costs by limiting the max_tokens parameter on each response so the model doesn't generate unnecessary text during iterative reasoning loops. Does that work for you?
Life saver! Semantic caching for agentic loops can slash costs dramatically—especially with the new Sol API’s fivefold output pricing gap, where each recursive reasoning step adds up fast. Have you tried implementing it?