Can LLMs grade their own homework to improve accuracy without manual effort?
Prompt engineering often forces a choice between two unsatisfying paths. Few-shot prompting demands hours of meticulous example crafting, yet even minor changes to those examples can destroy performance. Zero-shot prompting skips that work but relies on vague instructions like "think step-by-step," which may still produce confident but incorrect answers. One method is labor-intensive; the other is unreliable.
Consistency-based Self-adaptive Prompting (COSP) resolves this dilemma by letting the LLM generate and evaluate its own training examples. Instead of manually writing prompts, the model creates them, then verifies consistency to select the best ones for inclusion.
How the process works
The core insight is that output variance isn’t random noise—it reveals reliability. Asking a model the same question multiple times produces identical answers when correct and inconsistent ones when guessing. COSP exploits this by treating consistent model-generated responses as "pseudo-labels" to build few-shot prompts automatically.
Applying COSP in practice
Implementation follows a structured four-step loop:
- Generation Phase: Feed the model a batch of queries and produce multiple responses per query using higher temperature sampling.
- Filtering Phase: Identify responses matching the majority answer for each query, marking them as "high confidence."
- Adaptive Prompting Phase: Use those high-confidence question-answer pairs as dynamic in-context examples for remaining queries.
- Final Vote: Generate answers with the self-built prompt, then take the majority vote again.
Performance comparison
Standard approaches show clear trade-offs:
- Zero-shot requires minimal setup but suffers from high variance, often leading to irrelevant outputs.
- Manual few-shot delivers precision when examples are perfect, yet fails if any detail changes or scales poorly.
- COSP automates example selection, improving stability by adapting to the dataset’s specific patterns.
For production systems, this eliminates the guesswork of prompt tuning. Rather than manually selecting examples, the model identifies what it deems consistent, creating a self-correcting workflow that feels more dependable.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
This works wonders for coding tasks. By running the model's output through a consistency check—asking the same question multiple times and keeping only the answers that repeat—you get pseudo-labels that filter out the hallucinations every time. Self-correction catches those silly mistakes, and the consistency filter makes sure only the reliable ones stick.
I'm skeptical about this—does it handle complex logic or just the easy stuff? One concrete step is to ask the model the same reasoning question five times, then use four identical answers as a likely-correct pseudo-label in the prompt.

Adding a critique step before the final answer usually fixes those annoying logic gaps—by explicitly asking the model to evaluate its own reasoning for consistency, you can identify where it might be hallucinating or misapplying logic. This approach, like COSP, lets the model generate and verify its own high-quality examples dynamically, reducing the need for manual labor while ensuring reliability.