TeamStation AI / Research / Governance Research / Axiom Cortex Evidence Scoring Upgrade: No Proof, No Point
See how TeamStation tightened Axiom Cortex evidence scoring with strict proof, repeatability, fairness controls, and traceable results.
TeamStation tightened Axiom Cortex so every score begins with direct candidate evidence, survives repeatability checks, and remains traceable to the exact job criterion.
The easiest way to inflate an interview score is to let related language count as proof. A candidate says something technical nearby, the interviewer explains the missing part, an evaluator connects the dots, and the final number looks cleaner than the evidence.
We tightened Axiom Cortex engineer vetting around that failure.
The updated architecture follows one strict order: prove the evidence, lock the evidence, run the mathematics, then release a result only when the required controls agree. Missing proof receives no assumed credit. Interviewer comments cannot help the score. A broad answer cannot earn high marks only because it sounds technical.
No proof, no point. No agreement, no score.
Short answer
The Axiom Cortex evidence scoring upgrade makes each point traceable to four things: the job requirement, the exact question, the configured answer criteria, and the candidate's own words.
That makes scoring tougher, because related answers and evaluator assumptions no longer fill evidence gaps. It also makes the result easier to inspect, because a customer can move from the final score back to the criterion, transcript evidence, ownership decision, control status, and calculation receipt.
The release improves evidence accuracy, scoring consistency, and auditability. It does not claim that the system has already proven future job performance. Predictive validity requires longer term study against real work outcomes, and we are keeping that claim separate.
What accuracy means inside the release
Accuracy can become a fuzzy word fast, so we are defining it at the control level.
For the release, accuracy means:
- Job requirements are locked before candidate evidence enters the process.
- Must Have requirements become observable criteria.
- Each question is checked against its configured ideal answer criteria.
- The complete transcript is reviewed one question at a time.
- Direct candidate evidence must support every scoring result.
- Numeric calculations follow one fixed software policy.
- Repeated runs must agree before any score can be released.
- The accepted result can be traced back to its source evidence.
The control chain is closer to a chain of custody than a normal interview score. A lab sample without a label is not rescued because it resembles another sample. The evidence has to belong to the right job criterion, the right question, and the right speaker before it enters the scoring path.
The same boundary matters in our wider research. The constraint shift test asks whether engineering judgment survives a changed condition. The human task and agent alignment stress test asks which assumptions still require real human validation. The upgrade handles a different problem: whether the score can be reconstructed from admitted evidence without evaluator guesswork.
Seven controlled layers
Axiom Cortex now separates evidence interpretation from numeric scoring. The evidence layer can classify what the candidate demonstrated, but it cannot invent a score.
| TeamStation layer | Controlled responsibility |
|---|
| Role Blueprint Compiler | Converts the job description, Must Haves, questions, ideal answers, and follow ups into observable job criteria. |
| Evidence Mapping Engine | Reviews the complete transcript against every planned question and finds exact candidate evidence. |
| Evidence Decision Layer | Records opportunity, support, contradiction, missing proof, and ownership boundaries. |
| Deterministic Scoring Kernel | Applies the fixed scoring policy after the evidence record is locked. |
| Repeatability Gate | Requires five fresh evidence passes to agree after normalization before scoring continues. |
| Fairness Firewall | Blocks identity and language form from changing technical scores. |
| Result Lock | Prevents the same governed input from producing a different accepted result later. |
The Evidence Decision Layer is qualitative only. It cannot select or produce numeric scores, criterion weights, axis anchors, formula results, stars, hiring recommendations, or client submission decisions.
That separation matters because evidence interpretation and arithmetic are different jobs. Letting one flexible step do both creates room for scoring discretion to hide inside a confident explanation. Axiom Cortex now locks the evidence record before the numeric policy begins.
What counts as evidence
The evidence process performs six bounded actions:
1. Break each ideal answer into observable job criteria. 2. Find the smallest exact candidate quote that supports each criterion. 3. Confirm that the candidate had a fair opportunity to answer. 4. Separate direct technical support from contradiction or missing proof. 5. Separate personal ownership from shared team delivery. 6. Return one strict evidence record for deterministic scoring.
Every transcript segment must be assigned to one of four buckets:
- direct candidate evidence for the current question
- candidate evidence that belongs to another question
- interviewer or ghost evidence
- irrelevant material
Only direct candidate evidence from the same question can earn score credit. The other buckets remain useful for audit and context, but they cannot quietly become points.
The scoring path cannot award credit from interviewer explanations, interviewer praise, cross question evidence, unsupported assumptions, general topic similarity, ideal answer language not spoken by the candidate, resume claims not demonstrated during the interview, vague ownership claims, or answer length by itself.
That rule sounds severe until a buyer reviews the alternative. If an interviewer supplies the missing architecture, then praises the candidate for agreeing, the interview has measured cooperation with the interviewer, not independent technical evidence.
The public proof object
Customers do not need our private scoring formulas to inspect whether the process stayed governed. They need a receipt that shows the evidence path.
For every accepted criterion, the public facing proof object can show:
| Proof field | Buyer question |
|---|
| Role criterion | What exact job requirement was measured? |
| Question binding | Which planned question created the opportunity to demonstrate it? |
| Candidate evidence | What exact words support or contradict the criterion? |
| Ownership boundary | Did the candidate personally perform the work or describe team delivery? |
| Evidence status | Was the criterion supported, contradicted, or not demonstrated? |
| Repeatability status | Did the five fresh evidence passes agree after normalization? |
| Calculation receipt | Did the fixed scoring process return one identical accepted result? |
| Fairness status | Did the protected and proxy attribute controls pass? |
The proof object changes the buyer conversation. Instead of asking whether a score feels right, a CTO can ask which evidence earned each point, which requirement stayed unproven, and which control would have blocked the result.
The CTO proof system and nearshore engineering performance metrics extend that same operating idea after evaluation. Pre hire evidence should be inspectable, and delivery evidence should remain inspectable once the engineer enters the real system.
Repeatability without score shopping
Axiom Cortex does not run several evaluations and select the highest score.
Five fresh evidence passes must agree exactly after normalization. The process does not use majority voting, averaging, median selection, a preferred pass, four out of five approval, hidden retries, or manual score adjustment.
When a classification differs, one controlled review cycle may inspect only the disputed criteria. If disagreement remains, the evaluation stops. It does not pick the most favorable interpretation and keep moving.
After the evidence record is locked, the complete numeric calculation runs 100 times. All 100 calculations must return one identical result and one canonical hash. The 100 run control tests execution consistency under the fixed policy. It does not prove that the policy predicts future job performance.
That distinction is important. Repeatability asks whether the same governed input produces the same accepted output. Predictive validity asks whether the output relates to later real world outcomes. The first is a software and process control. The second requires a longer research program.
Why scores may move lower
Candidates did not suddenly become less capable. The measurement became more strict.
Scores may move lower because partial evidence does not become full support, vague evidence does not become technical proof, and missing evidence becomes not demonstrated. Explicit technical errors can cap an affected competency. Unsafe procedures can reduce a question score. Unsupported claims of personal ownership can reduce technical accuracy. Each question must stand on its own admitted evidence.
A high score now needs consistent evidence across procedural execution, mental model, technical accuracy, conceptual clarity, and organized retrieval. That is a harder bar, and it should be.
The neuro psychometric vetting model is useful only when its signals stay attached to the work. The Distributed Engineering Operating System then carries those decisions into team design, while the Nearshore Control Plane keeps delivery, device, access, cost, and governance evidence visible after launch.
Fairness is a gate, not a slogan
Technical scoring cannot change because of a candidate's name, country, nationality, culture, accent, first language, grammar, spelling, speaking speed, pauses, filler words, politeness, school prestige, interviewer praise, or native like wording.
The Fairness Firewall is designed to keep those identity and language signals outside technical scoring. If the fairness control fails, Axiom Cortex blocks the result and releases no score.
The Fairness Firewall is an engineering control, not proof that every population has already been validated. Our research on counterfactual bias testing for engineer scoring explains why protected attribute invariance and leakage testing are useful evidence, but neither one proves total fairness by itself. NIST's AI Risk Management Framework makes the wider point: validity, reliability, fairness, transparency, accountability, privacy, security, and safety have to be governed together.
What customers will receive
The customer view is designed to show the overall score, each question score, competency alignment, Must Have coverage, exact evidence, missing or contradicted requirements, ownership boundaries, score gates and caps, calculation receipts, fairness status, repeatability status, and a plain English explanation of why each point was earned or lost.
The goal is not a larger dashboard. The goal is a result that can survive a skeptical review.
TeamStation is building that evaluation layer inside one Distributed Engineering OS, because candidate evidence, team topology, delivery governance, and production telemetry should not live as disconnected opinions. LATAM is the application layer here. The evidence rule stays technical, while the operating system makes it usable across distributed teams.
Current release status
The upgraded architecture is in controlled validation before production activation.
The current implementation does not make hiring decisions, release production recommendations, or replace human review. Production activation remains locked until the remaining science, contract, security, representative validation, and release approvals are complete.
That boundary is deliberate. Stronger software controls can prove that the process executed consistently. Responsible production use also requires real world validation, fairness review, security review, accommodations, representative samples, and measured outcomes.
We made the score harder because the proof should carry the weight, not the confidence of the evaluator. Evidence comes first. Mathematics comes after the lock. A result leaves the system only when every required control passes.
No proof, no point.
No agreement, no score.
Related TeamStation research
Sources, authorship, and limitations
Lonnie McRorey directed the release argument, evidence boundaries, scoring controls, and TeamStation application. The public article explains the operating method without publishing private formulas, weights, candidate records, interview transcripts, or protected data.
The article does not claim that Axiom Cortex is a clinical instrument, that the upgrade guarantees fair outcomes, that repeatability proves predictive validity, or that one score should decide whether a person is hired. The architecture remains under controlled validation, and human review remains required.