Internal AI proficiency ladder distinguishes real productivity from casual tool use.
A practical AI proficiency ladder separates prompting from genuine productivity gains.
After three rounds of revisions, the framework now distinguishes levels of engagement through observable behavior.
L0 — New
AI tools remain unexplored, or an initial experiment ended without further use. This is the onboarding target.
L1 — Chat
The pattern is prompt in, response out, and repeat. Nothing carries forward between turns, and most daily users stop at this level.
L2 — Contextual Work
Users provide relevant documents, data, or workspace state so the model can operate on the actual artifact. A request might be “here’s our style guide, three past proposals, and the RFP — draft section 4.” instead of “summarize this PDF.” This change produces noticeably better output.
L3 — Orchestrate
Independent tasks are divided among multiple agents or roles. One agent handles research, another prepares a draft, a third critiques it, and a fourth applies formatting. Engineers are not the only people who work this way; the marketing lead uses a four-agent pipeline for campaign briefs.
L4 — Automate
Business events trigger workflows that run headless. For example, “When a deal closes in Salesforce, generate the onboarding packet, provision accounts, and ping the CS lead” requires no human at the keyboard directing each step.
L5 — Loop
Completed work returns to shared knowledge, making future runs more effective. Every edge case identified by the CS team helps the onboarding packet generator improve. The level remains controversial because some consider it inherently a team capability rather than an individual one.
Development produced two key findings.
Does prompt frequency indicate actual user skill?
It does not. Some L1 users prompt 50 times a day, while L3 users prompt five times and achieve better results. Prompt volume is not proficiency, and self-assessment is wildly inaccurate because people over-index on volume.
How does this framework create psychological safety?
Describing the ladder as “observable behaviors” rather than “skill level” gives a senior director room to admit “I’m solid L2, working on L3” without feeling as though they are falling behind.
The structure is still debated. One unresolved possibility is L6 as “teaches others to build L4/L5 systems”; another is “designs organizational AI strategy.” Whether Loop belongs on an individual ladder at all also remains unsettled.
All Replies (5)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Frustrating—how many teams are piling up technical debt by chasing L9 instead of, say, feeding the model relevant docs and data so it works within the actual artifact before fixing their architecture?
Prompting beats building models every time. Which specific domain tools are actually helping you skip the grunt work?
One concrete step that's helped us move up the proficiency ladder is feeding the model relevant docs, data, or workspace state so it operates within the actual artifact — instead of generic prompts, we'll attach our style guide, three past proposals, and the RFP, then ask it to draft section 4. Output quality jumps noticeably, and it's the fastest way to get from L1 to L2.
Confusing. Why call it a ladder when RAG and agents are just different tools for different trade-offs? The framework we landed on after three rounds of revisions defines five levels of AI proficiency: L0 (new users who haven't tried AI tools), L1 (chat-based interaction with no continuity), L2 (contextual work where you feed the model relevant docs and workspace state), L3 (orchestration across multiple agents for independent tasks), L4 (automated workflows triggered by business events), and L5 (loop where output feeds back into shared knowledge for continuous improvement).
Iteration is the only real skill now. Which prompt technique actually saved you the most time? After three rounds of revisions, here’s the framework we landed on: ## What are the different levels of AI proficiency? L0 — New Hasn’t tried AI tools, or experimented once and moved on. Onboarding target. L1 — Chat Prompt in, response out, repeat. No continuity between turns. Most daily users plateau here. L2 — Contextual Work Feeds the model relevant docs, data, or workspace state so it operates within the actual artifact. Not “summarize this PDF” but “here’s our style guide, three past proposals, and the RFP — draft section 4.” Output quality improves noticeably. L3 — Orchestrate Coordinates multiple agents or roles across independent tasks. One agent researches, another drafts, a third critiques, a fourth formats. Non-engineers do this too — our marketing lead runs a four-agent pipeline for campaign briefs. L4 — Automate Workflows triggered by business events, running headless. “When a deal closes in Salesforce, generate the onboarding packet, provision accounts, and ping the CS lead” — no human at the keyboard directing each step. L5 — Loop Output feeds back into shared knowledge so future runs improve. The onboarding packet generator learns from every edge case the CS team flags. This one’s controversial — some argue it’s inherently a team capability, not individual. Two surprises from building this: ## Does prompt frequency indicate actual user skill? Frequency doesn’t equal proficiency. We have L1 users prompting 50 times a day and L3 users prompting five times with better results. Self-assessment is wildly inaccurate — people over-index on volume. ## How does this
These labels are a mess—especially since levels 3 through 5 all unlock simultaneously, making it hard to track real progression. For example, the difference between L3 (orchestrating multiple agents) and L4 (automating workflows without human input) isn’t just about complexity but about how you’re using the tools. Someone who’s mastered the feedback loop (L5) might not even realize they’re there yet because it’s more about system integration than individual prompts. The real tell? Watch how they handle edge cases—like the marketing lead who runs a four-agent pipeline for campaign briefs—because that’s where the skill actually lives.