Operator Engineer evidence map

Proof you can recognize.

Badges are how the AI Work Assessment names what your work shows. Each one is a lens into cited work, and the proof behind it travels with it.

Fourteen badges, no ladder People can earn several, and no badge outranks another.
Observed ★★ Established ★★★ Demonstrated Stars show proof strength, not skill level.
Why this exists

AI capability is hard to see from conventional signals.

Resumes, titles, and portfolios fail to show how someone actually works with AI. Overflow looks for observable decisions, repeated behaviors, and outcomes across real work.

Real work over self-report

Work arcs and outcomes carry more weight than claimed expertise, job titles, or a list of installed tools.

Behavior over volume

Activity and cadence give context; the evaluator looks for what the person chose, changed, checked, and shipped.

Evidence over labels

Every badge points back to badge-specific proof, and a missing badge means the evidence did not clear its gate.

Visible evidence only

The evaluator never reconstructs private chain-of-thought. It inspects visible requests, revisions, tool use, verification, and recovery, separates the person's choices from an agent's mistakes, and labels every coverage gap.

How the assessment works

Private by default. Reviewable before sharing.

Raw histories stay local. Only the finished, participant-reviewed profile is uploaded, and only if the participant opts in.

1 · Choose sources

The participant selects agent histories, GitHub activity, and optional career context, each with a dated coverage window.

2 · Extract locally

The evaluator reads retained evidence on the participant's machine and separates observed work from self-reported context.

3 · Build work arcs

Related sessions become real work arcs with a problem, decisions, artifacts, outcomes, dates, and confidence level.

4 · Apply badge gates

All 14 badges begin at zero. Qualifying proof, independent observations, outcomes, time span, and counterevidence determine each result.

5 · Show the whole map

Awarded badges, unawarded badges, coverage limits, and the proof needed for the next star are shown together.

6 · Owner review

The participant checks sensitive abstractions and factual claims, then decides whether to keep the HTML local or publish a profile.

How to read a result

What a badge says, and what it does not.

A badge is a compact index into evidence. It is useful only when its definition, proof, limits, and counterevidence travel with it.

What a badge says

  • A defined behavior was observed, clearing the badge's positive gate.
  • The proof strength is stated. Stars describe breadth, independence, outcomes, and time span.
  • The claim is inspectable. A reader can follow it back to specific work arcs and the closest misses.
  • The result is bounded by the source window.

What a badge does not say

  • Not an intelligence, personality, or seniority score. Confidence, polish, verbosity, and fluency are not proxies.
  • Not a hiring decision or certification. It starts a conversation about demonstrated work.
  • Not a complete career history. Local tools retain different amounts, and missing evidence is not negative evidence.
  • Not a ranking of worth or potential. Stars are never combined into one score.

Technical chops

How the person builds.

Technical chops★★T-01

Prototyper

Technical chops★★T-02

FrontendCrafter

Technical chops★★★T-03

ProductionShipper

Technical chops★★★T-04

SystemsArchitect

Technical chops★★T-05

ContextEngineer

Technical chops★★★T-06

AgentOrchestrator

Business know-how

What the person understands and translates, including industry and functional domain.

Business know-how★★★B-01

WorkflowArchitect

Business know-how★★B-02

ValueTranslator

Business know-how★★★B-03

ClearCommunicator

Business know-how★★B-04

AdoptionOperator

Good judgment

How the person decides, verifies, recovers, and maintains systems over time. Each one requires observable decisions.

Good judgment★★★J-01

VerificationFirst

Good judgment★★J-02

TradeoffNavigator

Good judgment★★J-03

RecoveryOperator

Good judgment★★★J-04

SystemsSteward

Proof strength

What one through three stars means

Stars describe the breadth, independence, and confidence of the evidence for one badge. They are never added into a person score.

Observed

One concrete badge-specific episode, or several observations inside one narrow arc. It proves the behavior happened.

★★

Established

At least two distinct qualifying arcs, or one sustained 30+ day arc with three independent observations; at least one direct outcome and no unresolved material contradiction.

★★★

Demonstrated

At least three badge-specific arcs across two systems or organizations and 90+ days, with two high-confidence arcs and two direct outcomes.

Anti-inflation rules

How we keep badges honest

The assessment starts at zero for all fourteen badges and has to earn its way to each one.

Every badge gets a verdict

All fourteen are rated or explicitly not awarded. A whole pillar may come back empty.

Evidence is spent once

One piece of evidence supports one badge. The same work arc counts twice only when it shows two genuinely different decisions or outcomes.

The near misses are shown

Contradictions and closest misses are kept. Every one- or two-star badge names the evidence that would earn the next star.

Badge by badge

What each badge requires

Every badge has a positive gate and a specific “not enough” rule.

Prototyper

Qualifies: repeatedly turns ambiguous ideas into functional artifacts a real user can try; two working-prototype-or-stronger arcs.

Not enough: mockups, plans, isolated experiments, or tool usage without a usable artifact.

Frontend Crafter

Qualifies: recurring visual, interaction, responsive, accessibility, or design-system judgment on interface work.

Not enough: React/CSS presence, hosted builders, generated imagery, or a polished screenshot without demonstrated decisions.

Production Shipper

Qualifies: repeatedly gets substantive changes into real use with merge, release, deployment, migration, or handoff evidence.

Not enough: local apps, staging links, prototypes, plans, or “ready to ship” claims.

Systems Architect

Qualifies: repeated consequential decisions about technical boundaries, data models, dependencies, integrations, migrations, or reliability that survive implementation.

Not enough: general systems language, broad ownership, diagrams, or technology selection without a consequential design decision.

Context Engineer

Qualifies: repeatedly structures, retrieves, refreshes, compresses, or hands off durable agent context through reusable mechanisms.

Not enough: long prompts, a large context window, copied background, or an installed memory tool.

Agent Orchestrator

Qualifies: decomposes work across agents, scopes ownership, integrates results, and verifies the combined outcome.

Not enough: spawning subagents, parallel calls, or token volume without synthesis and control.

Workflow Architect

Qualifies: maps actors, decisions, exceptions, controls, and data movement into a scalable business workflow or automation.

Not enough: generic automation, a connector, or a linear happy-path script without real process understanding.

Value Translator

Qualifies: repeatedly connects goals, economics, KPIs, customer needs, or operating constraints to technical scope and priority.

Not enough: requirements transcription, strategy language, or stakeholder contact without an observable outcome link or tradeoff.

Clear Communicator

Qualifies: across three substantive tasks, supplies a usable outcome, relevant context, constraints, and definition of done; corrections are specific and exploration closes in a decision.

Not enough: politeness, confidence, verbosity, grammar, native fluency, long prompts, agent agreement, or one unusually good brief.

Adoption Operator

Qualifies: repeatedly designs handoff, training, governance, review cadence, change management, or the operating loop after delivery.

Not enough: a demo, documentation alone, launch messaging, or a one-time handoff.

Verification-First

Qualifies: repeatedly uses independent checks before acceptance: tests, visual QA, source comparison, reconciliation, production checks, or second-agent audit.

Not enough: the producing agent saying it works, casual review, or one check performed only after failure.

Tradeoff Navigator

Qualifies: repeatedly chooses among real alternatives using value, cost, latency, risk, maintainability, reversibility, or time.

Not enough: preferences, tool loyalty, changing tactics without new evidence, or a decision with no stated consequence.

Recovery Operator

Qualifies: detects a real failure, diagnoses it, constrains retry or rollback, restores safe state, and changes the approach.

Not enough: ordinary iteration, repeated retries, or a successful first attempt.

Systems Steward

Qualifies: improves a live or established system over time through monitoring, maintenance, refactoring, reliability, or governance.

Not enough: shipping the initial release, claiming ownership, or returning once for an isolated fix.

System rule

Every badge has a separate evidence test. Prototyper is not a junior Production Shipper. Systems Architect is technical structure; Workflow Architect is business-process structure. Production Shipper gets work live; Systems Steward proves post-launch care.