Framework · Ethotechnics

How do you know what you know — and what happens when you're wrong?

An open framework that makes consequential AI claims verifiable and correctable by treating evidence, authorization, challenge, reconsideration, correction, and escalation as explicit system states.

It is not a separate body of work from the products. It is why they are shaped the way they are: NextConsensus holds the evidence state, Ambit the authorization state, Refract the change state. Each one makes explicit a state that consequential systems usually leave implicit.

Every AI system deployed in high-stakes environments makes implicit claims: that a guideline is current, a drug protocol is safe, or a patient needs immediate prioritization.

A claim is verifiable when you can see what backs it, challenge is possible, and corrections reach the systems that relied on it.

Most AI governance work starts with the operating model — who approves what, under what conditions, for how long. That is downstream. First: how do you know what you know, and how is an error repaired?

NextConsensus freezes each forecast before the outcome is known. Refract keeps source changes verifiable, and Fast Harm, Slow Repair applies the same discipline to recovery evaluation.

Ambit implements the operating model in agent infrastructure: authorization held apart from capability, changes with their own inverse, verification before trust, and bounded tool scope.

Ethotechnics publishes numbered, versioned standards. Each is crosswalked to NIST AI RMF, ISO/IEC 42001, and the EU AI Act; the mapping is mine, not an endorsement by those bodies. No outside institution has adopted one yet, so they remain proposals.
ID Title Version Status NIST AI RMF ISO/IEC 42001 EU AI Act
STD-01 The Temporal Bill of Rights v1.0 Draft
STD-02 Contestability & Recourse v1.1 Draft
STD-06 Human Impact Safety Case v0.5 Draft
STD-07 Revisable Delegation Record Conformance checker → v0.1 Draft
STD-08 Delegation v0.2 Draft
+4 Indexed, not yet readable Justice SLOs, verifiable-credential schemas, a minimum viable contestability standard, and an institutional failure postmortem template. They are named so the count is checkable; links appear when the documents resolve. A FHIR profile set was drafted and is now marked deprecated upstream, so it is not counted here. Indexed, not yet readable
  • The twelve laws Each law carries the invariant a standard clause binds: capability does not imply authority, authority decays unless renewed, every delegation creates a correction obligation, and nine more.
  • The Ethotechnical invariant One sentence the laws compress to: no system may accumulate consequential agency faster than the institution accumulates the capacity to inspect, challenge, revise, and survive its decisions.
  • Core axioms Five commitments the laws are owed to, with the argument for why they are entitlements rather than preferences.
  • Theory Essays on why the laws hold. Motivation only: no standard cites one as a requirement, and a reader can adopt a clause without the argument.

What I require from the teams I lead. Each one links to the work it came out of.

  1. 01 Human override is a product feature.

    Override, escalation, and return-for-review are core capabilities that determine whether a high-stakes system can safely deploy.

    Recurs in two deployments

    • Andwise Automated clause analysis surfaced and explained each issue. An accountable human reviewed every analysis before a physician could act on it.
    • Ethotechnics Authorization states written down as an open standard: who may override, and what the override obliges them to record.
  2. 02 Provenance beats confidence scores.

    Knowing where a recommendation came from — and what evidence backed it — matters far more than an opaque confidence score nobody can audit.

    Recurs in two deployments

    • Refract Every change replayed into a verifiable event carrying its provenance, with the judgment left to the caller.
    • NextConsensus A public ledger of forecasts registered and frozen before the outcome is known, scored against the record afterwards.
  3. 03 The decision the AI feeds into matters more than the AI's output.

    When an automated system causes harm, the decisive question is never just 'was the model wrong?' It is 'who owned the decision the model fed into?'

    Recurs in two deployments

    • Epic Installed, signed off, and live — and still not doing the job, because nobody owned what happened after the go-live.
    • The Crumple Zone Essays on the gap between an automated recommendation and the person who has to carry it out.
  4. 04 Governance belongs in the product.

    Evaluation, monitoring, escalation, and correction work best when designed directly into daily product workflows rather than managed via external committees.

    Recurs in two deployments

    • Ethotechnics Authorization, correction, and escalation published as open, versioned standards.
    • Andwise Compliance review was a routed step in the flow with accountable sign-off.
  5. 05 A deployment is finished when corrections reach every record it touched.

    An error corrected at the source remains active everywhere it already propagated.

    Recurs in two deployments

    • Refract Downstream systems receive what changed and when.
    • Fast Harm, Slow Repair A protocol that measures how far a wrong output travels before the correction catches up with it.
  6. 06 Evaluate workflows before models.

    A model that performs well in evaluation can still fail at the integration point where busy clinicians have to use it.

    One instance so far

    • Transcarent Four programs shipped on one shared decision architecture because routing was the binding constraint.
  7. 07 Missing approval can be earned back; a harmed patient can't.

    A premature launch creates downstream costs that dwarf any speed advantage when the risk is clinical.

    One instance so far

    • Epic The escalation route that had to exist before the software could honestly be called live.
  8. 08 Hospital IT trusts the system around the model.

    Hospital IT, security, and legal teams evaluate the whole operational system surrounding a model — reviewing access boundaries and audit trails as closely as the interface.

    One instance so far

    • Doximity Hospital security reviews cleared the product on documented policies and procedures.

AI deployment requires continuing authorization: when a system may influence decisions, what conditions limit that authority, how people can challenge it, and what evidence or behavior reopens review.

A model recommends a treatment change. Authorization stays versioned and conditional the whole way.

State transitions for the treatment-change recommendation. Rows are in order; a state is entered only when the row above it exits.
State Trigger Who acts Evidence gate What propagates
State 0 Not authorized A use is proposed. Named deployment and escalation owners
  • Defined clinical or operational decision
  • Named deployment and escalation owners
  • Specified affected population and exclusions
Nothing. The proposed use cannot influence care or workflow.
State 1 Observed The recommendation is observed but cannot influence care. State 0 exits when A bounded use case, accountable owner, and evaluation plan are approved.
  • Silent or retrospective evaluation
  • Error taxonomy and exception review
  • Baseline comparison against current practice
Nothing to care or workflow. Outputs are captured for evaluation.
State 2 Constrained use After a silent evaluation, it may inform a narrow workflow under human review. State 1 exits when Observed performance and failure modes justify a limited prospective deployment. Explicit human review on each recommendation
  • Prospective workflow validation
  • Documented override and escalation paths
  • Monitored safety, equity, and operational indicators
Recommendations into one narrow workflow, within limited population scope and predefined stop conditions.
State 3 Routine reliance State 2 exits when The deployment performs acceptably inside its stated scope and the institution can pause, correct, or roll it back. Independent review, with challenge rights preserved
  • Stable prospective performance
  • Independent evaluation appropriate to the use
  • Operational readiness for correction and rollback
Routine reliance within the defined scope, with change detection and periodic reconsideration.
Any state Reopened A new safety signal suspends it immediately. 01 Detect Identify a potentially material change in evidence, model behavior, policy, data, workflow, population, or observed outcomes. The clinical safety lead opens the reconsideration case and issues the disposition. 04 Review → Reconsideration case Present the source-traced change, prior rationale, observed performance, dissent, and unresolved uncertainty to the authorized reviewers. 06 Propagate → Corrected operating state Update the authorization record and each connected workflow, instruction, interface, monitoring rule, and affected stakeholder.

The loop opens on a change, not on a review cycle. A system that can only be reconsidered on schedule is unrevisable in between.

Explainable

A clinician, compliance officer, or patient can see what the system considered and recommended, with enough detail to understand the basis for action.

Challengeable

The recommendation can be overridden, escalated, or sent back for review without halting care or creating a compliance incident.

Correctable

When the evidence, policy, or model changes, the decision can be updated and the record corrected. Someone is responsible for that correction.

Terms I use for AI deployment, institutional decision-making, and the boundary between human judgment and automated systems.
Term In a deployment
The property What is at stake.
Revisability Ensures automated decisions have explicit correction paths so mistakes don't become permanent policy.
Constraint Realism Toothless Ethics Evaluates safety claims by checking which unsafe or profitable actions the architecture prevents.
Failure modes Ways the property is lost.
Care Subsidy Toothless Ethics Eliminates reliance on human heroism by enforcing programmatic safety boundaries and automated escalation.
Governance by Attrition Pending Prevents systems from hiding behind delays — requires explicit, auditable approval or rejection.
Asymmetric Irreversibility If Every User Is a Potential Threat Balances automated decision speed with programmatic exoneration and rapid appeal.
Error Survivability You Don't Have the Right Measures how far model mistakes propagate across tools to determine where human review and rollback are needed.
Mechanisms Structures that keep it.
Review Trigger The Official Record Is Late Requires the firing conditions to be fixed before launch.
Reliance Memory The Official Record Is Late Dependent users and workflows are notified when source facts or model assumptions shift.
Decision Maintenance Provides structured ownership and review intervals for deployed AI models.
Evidence Dependency Graph Has to capture the dependency when the conclusion is drawn; after a retraction, tracing it is manual.
Operating condition What the mechanisms have to hold under.
Cognitive Scarcity How to Design for Cognitive Scarcity Requires that override, challenge, and review interfaces work reliably under stress and fatigue.

One page, ending in either a decision you can rely on or a list of what still needs systems work.

Open the checklist →
Bi-directional Trajectory

Accountable Routing & Escalation

First observed 2012

Designing explicit, programmatically enforced escalation paths so clinical and operational errors never stall without a responsible owner.

2012 Epic

Co-led CMS PQRS and built its federal quality escalation path, unblocking stalled reporting. View record →

2021 Transcarent

Architected unified cross-specialty exception queues and role-based clinical routing. View record →

2023 Essay

Formalized structural escalation in 'Toothless Ethics' — why principles fail without programmatic boundaries. View record →

2025 Ethotechnics & Ambit

Specified the Reconsideration Loop state machine and computable agent capability bounds. View record →

Trust as a Product Decision

First observed 2014

Proving identity and safety boundaries at the systemic level to eliminate downstream adoption friction.

2014 Doximity

Surveyed and ID-verified 35,000+ physicians by license, DEA number, and hospital domain. View record →

2016 Doximity Dialer

Proved that verified account identity enables unverified caller ID routing, passing hospital IT review. View record →

2025 Ambit

Separated capability from authorization in agent infrastructure — verifying what a tool may do before execution. View record →

Revisability & Error Propagation

First observed 2011

Ensuring that when evidence, policy, or models change, relying systems automatically update and correct affected downstream workflows.

2011 Georgia Tech RNA

Modeled thermodynamics and barrier kinetics deciding whether a molecular process completes or stalls. View record →

2024 Refract & NextConsensus

Built deterministic change-detection bots and claim-trajectory scoring against public records. View record →

2025 Fast Harm, Slow Repair

Designed a recovery-evaluation scaffold measuring harm duration and correction propagation; the draft covers one of twelve development cases. View record →

Cognitive Scarcity & User Navigation

First observed 2019

Designing interfaces and decision workflows for real people operating under severe stress, distraction, and fatigue.

2019 CancerCompass

Restructured oncology navigation: replacing dense content with immediate actionable sequencing. View record →

2022 Essay

Published 'How to Design for Cognitive Scarcity' — designing systems that don't assume hero users. View record →

2023 Andwise

Designed physician employment contract analysis to highlight hidden restrictive covenants instantly. View record →

I study what changes when a decision that used to end with a person ends in software instead: where errors travel, who absorbs the work they create, and whether the system can repair itself when the information under it changes.

Three cases, and which of the questions each one asks
Case Mechanism Who absorbs the cost Which questions
Emergency department loudspeaker Patients are called for triage by name over a loudspeaker. For a Deaf patient the mechanism deciding who waits is inaudible. The Deaf patient; what it costs is measured in mortality.
  • Who waits?
  • What cannot be allowed to fail?
Epic: a misrouted clinical alert A misrouted alert could bury a critical lab result. Whichever clinician trusted the queue. The routing system carried no matching accountability.
  • Who absorbs error?
  • What gets buffered?
Andwise: monetization The monetization paths most likely to fund growth would have made employers or financial institutions the customer. The physician, the party the product answered to — preserved by shutting the company down.
  • What is expendable?
  • What cannot be allowed to fail?

In 2012 I helped build a medication-recommendation system for type 2 diabetes, and we measured the thing that now worries me most about deployed models: not whether the recommendation was right, but what happened to the clinician's own judgment once they had seen it.

  • Internists n=2 62% 92% change +30 pts
  • Familiar with the rules n=4 64% 86% change +23 pts
  • Endocrinologists n=4 68% 76% change +8 pts
  • Unfamiliar with the rules n=2 71% 71% change +1 pts
Agreement with the algorithm, before and after the clinician saw its recommendation. Blinded validation, twenty patient data sets per reviewer, six reviewers in total. Percentages are rounded; the subgroups nest rather than partition, so the reviewer counts do not sum to six. Rows are ordered by how far each group moved.

The specialists barely moved and the generalists moved most. The group that moved least of all was the one that did not know how the algorithm worked, and the group that understood its rules moved almost as far as the generalists. Understanding the tool predicted deferring to it.

A recommendation moves the judgment beside it, and it moves different readers by different amounts. This course project's data shows the direction but cannot size the effect. That is the mechanism the authority-boundary work exists to constrain: a model's output becoming institutional policy without an explicit decision.

Observed evidence what the sources actually show, timestamped and re-checkable Assessment the model's read on whether the evidence has moved — and by how much Institutional authority automatic AUTHORITY BOUNDARY — crossed only by a named person a decision, on the record Preserve Caveat Narrow Expand Escalate Retire Monitor / defer No action
The constraint the measurement above argues for. Evidence may move an assessment on its own; nothing moves an institutional position without a person crossing the boundary and leaving a record of having done it. “No action” is one of the dispositions, because a system where declining to act is not a recordable choice will drift into acting by default.

Jain, K., Patel, K., Rowland, J., Yong, C. DiaMonD (DIAbetes MONitoring and Dosing) System. Wallace H. Coulter Department of Biomedical Engineering, Georgia Tech / Emory University. Advised by Dr. Lawrence Phillips, MD.

  • Case studies Where this came from: five deployments through hospital security review, federal quality reporting, and clinical sign-off.
  • Decision records Four decisions with the reasoning as I wrote it down before the outcome was known, and the contemporaneous document attached where one survives.
  • Principles What 14 years of clinical products left me unwilling to ship without.
  • Reliance Lab One composite deployment, six months in, with new information on the table. Make the call and see what it commits the organization to.