Real work over self-report
Work arcs and outcomes carry more weight than claimed expertise, job titles, or a list of installed tools.
Badges are how the AI Work Assessment names what your work shows. Each one is a lens into cited work, and the proof behind it travels with it.
Resumes, titles, and portfolios fail to show how someone actually works with AI. Overflow looks for observable decisions, repeated behaviors, and outcomes across real work.
Work arcs and outcomes carry more weight than claimed expertise, job titles, or a list of installed tools.
Activity and cadence give context; the evaluator looks for what the person chose, changed, checked, and shipped.
Every badge points back to badge-specific proof, and a missing badge means the evidence did not clear its gate.
The evaluator never reconstructs private chain-of-thought. It inspects visible requests, revisions, tool use, verification, and recovery, separates the person's choices from an agent's mistakes, and labels every coverage gap.
Raw histories stay local. Only the finished, participant-reviewed profile is uploaded, and only if the participant opts in.
The participant selects agent histories, GitHub activity, and optional career context, each with a dated coverage window.
The evaluator reads retained evidence on the participant's machine and separates observed work from self-reported context.
Related sessions become real work arcs with a problem, decisions, artifacts, outcomes, dates, and confidence level.
All 14 badges begin at zero. Qualifying proof, independent observations, outcomes, time span, and counterevidence determine each result.
Awarded badges, unawarded badges, coverage limits, and the proof needed for the next star are shown together.
The participant checks sensitive abstractions and factual claims, then decides whether to keep the HTML local or publish a profile.
A badge is a compact index into evidence. It is useful only when its definition, proof, limits, and counterevidence travel with it.
How the person builds.
What the person understands and translates, including industry and functional domain.
How the person decides, verifies, recovers, and maintains systems over time. Each one requires observable decisions.
Stars describe the breadth, independence, and confidence of the evidence for one badge. They are never added into a person score.
One concrete badge-specific episode, or several observations inside one narrow arc. It proves the behavior happened.
At least two distinct qualifying arcs, or one sustained 30+ day arc with three independent observations; at least one direct outcome and no unresolved material contradiction.
At least three badge-specific arcs across two systems or organizations and 90+ days, with two high-confidence arcs and two direct outcomes.
The assessment starts at zero for all fourteen badges and has to earn its way to each one.
All fourteen are rated or explicitly not awarded. A whole pillar may come back empty.
One piece of evidence supports one badge. The same work arc counts twice only when it shows two genuinely different decisions or outcomes.
Contradictions and closest misses are kept. Every one- or two-star badge names the evidence that would earn the next star.
Every badge has a positive gate and a specific “not enough” rule.
Qualifies: repeatedly turns ambiguous ideas into functional artifacts a real user can try; two working-prototype-or-stronger arcs.
Not enough: mockups, plans, isolated experiments, or tool usage without a usable artifact.
Qualifies: recurring visual, interaction, responsive, accessibility, or design-system judgment on interface work.
Not enough: React/CSS presence, hosted builders, generated imagery, or a polished screenshot without demonstrated decisions.
Qualifies: repeatedly gets substantive changes into real use with merge, release, deployment, migration, or handoff evidence.
Not enough: local apps, staging links, prototypes, plans, or “ready to ship” claims.
Qualifies: repeated consequential decisions about technical boundaries, data models, dependencies, integrations, migrations, or reliability that survive implementation.
Not enough: general systems language, broad ownership, diagrams, or technology selection without a consequential design decision.
Qualifies: repeatedly structures, retrieves, refreshes, compresses, or hands off durable agent context through reusable mechanisms.
Not enough: long prompts, a large context window, copied background, or an installed memory tool.
Qualifies: decomposes work across agents, scopes ownership, integrates results, and verifies the combined outcome.
Not enough: spawning subagents, parallel calls, or token volume without synthesis and control.
Qualifies: maps actors, decisions, exceptions, controls, and data movement into a scalable business workflow or automation.
Not enough: generic automation, a connector, or a linear happy-path script without real process understanding.
Qualifies: repeatedly connects goals, economics, KPIs, customer needs, or operating constraints to technical scope and priority.
Not enough: requirements transcription, strategy language, or stakeholder contact without an observable outcome link or tradeoff.
Qualifies: across three substantive tasks, supplies a usable outcome, relevant context, constraints, and definition of done; corrections are specific and exploration closes in a decision.
Not enough: politeness, confidence, verbosity, grammar, native fluency, long prompts, agent agreement, or one unusually good brief.
Qualifies: repeatedly designs handoff, training, governance, review cadence, change management, or the operating loop after delivery.
Not enough: a demo, documentation alone, launch messaging, or a one-time handoff.
Qualifies: repeatedly uses independent checks before acceptance: tests, visual QA, source comparison, reconciliation, production checks, or second-agent audit.
Not enough: the producing agent saying it works, casual review, or one check performed only after failure.
Qualifies: repeatedly chooses among real alternatives using value, cost, latency, risk, maintainability, reversibility, or time.
Not enough: preferences, tool loyalty, changing tactics without new evidence, or a decision with no stated consequence.
Qualifies: detects a real failure, diagnoses it, constrains retry or rollback, restores safe state, and changes the approach.
Not enough: ordinary iteration, repeated retries, or a successful first attempt.
Qualifies: improves a live or established system over time through monitoring, maintenance, refactoring, reliability, or governance.
Not enough: shipping the initial release, claiming ownership, or returning once for an isolated fix.
Every badge has a separate evidence test. Prototyper is not a junior Production Shipper. Systems Architect is technical structure; Workflow Architect is business-process structure. Production Shipper gets work live; Systems Steward proves post-launch care.