AI safety and accountability

How I approach AI accountability in production

A product leader’s view of safety architecture: evidence traceability, authorization design, review, escalation, and correction — grounded in shipped healthcare products.

Safety and deployment systems

Authorization states, reconsideration loops, and escalation designs intended for consequential AI deployments.

Judgment Commitment →

DeploymentGovernance

A proposed protocol for recording when an AI system's conclusions are reliable enough to act on. Its specification covers proposition registration, evidence-state freezing, and revision obligations.

Problem
AI systems make probabilistic claims. Institutions need deterministic commitments they can rely on and later revise.
Mechanism
Frozen evidence-state registration with explicit scope, authorization conditions, and revision obligations.
Potential application
Could support system cards, deployment safety cases, and other records that must connect evidence to authorized use.

Authorization states →

GovernancePolicy

A graduated permission model for AI deployment — from prohibited through constrained use to routine reliance. Each transition requires explicit evidence and institutional sign-off.

Problem
Binary approval/rejection fails for AI systems whose risks emerge over time.
Mechanism
Four states: Not authorized, Observed, Constrained use, Routine reliance — each with specific conditions.
Design intent
A staged authorization model for increasing reliance only as evidence, monitoring, and correction capacity mature.

Reconsideration loop →

OperationsSafety

Six-step protocol for responding when new evidence, performance degradation, or escalation triggers call a prior authorization into question.

Problem
AI systems operate continuously but the evidence they rely on does not stay current.
Mechanism
Detect → Connect → Triage → Review → Decide → Propagate. Each step has defined inputs, actors, and outputs.
Frontier relevance
A reusable pattern for maintaining safety across model updates, context shifts, and new-use-case expansion.

Escalation design →

SafetyHuman oversight

Five conditions that force human review: patient-safety signal, performance degradation, new evidence, repeated overrides, and model/prompt/data changes.

Problem
Automatic escalation is either too sensitive (desensitizes reviewers) or too permissive (misses failures).
Mechanism
Trigger conditions paired with response levels — advisory notice, mandatory review, or automatic override.
Frontier relevance
Scales to any domain where model output must be checked before consequential action.

Technical systems

Open-source systems and prototypes for evidence verification, capability tracking, and forecasting.

NextConsensus →

ForecastingEvidenceArchitecture

Forecasting platform and accountability protocol under development. It is designed to register frozen-evidence probability forecasts for observable medical-guideline, regulatory, and coverage transitions.

Protocol design
Point-in-time evidence state, immutable proposition registration, adjudication rules, probability history with Brier-score resolution.
Non-AI component
Deterministic observation layer (Refract) ingests structured change events from public sources. No model involved — pure verification and provenance.
Relevance
Separating observation from inference could also inform monitoring designs for model behavior in production.
S V

Refract →

Open sourceVerificationDeterministic

Open-source deterministic verification engine that checks whether a claim still holds against its sources. Produces structured change events with full provenance records.

Architecture
Provenance graph, source registration, claim-checking pipeline, structured event output. No model — pure observation.
Relevance
The infrastructure layer AI systems need to know when underlying facts change. Directly applicable to retrieval-augmented generation freshness monitoring.

Capability Graph →

Open sourceAgentsReliability

Prototype for tracking what an AI agent can do, what may be decaying, and the dependencies behind new capabilities. Built with Claude Code.

Architecture
Maturity scores, decay rates, bottleneck analysis, unlock paths. Directed graph of capabilities with edge weights for cost of acquisition.
Potential application
The graph structure could support longitudinal capability reviews; production drift detection has not yet been demonstrated.

Operational evidence

Shipped production systems across healthcare — a domain where the cost of automated error is measured in patient outcomes.

Epic EHR — Implementation Engineer

EHRDecision supportPolicy

Configured clinical decision support and quality-measurement workflows within Epic EHR for federal incentive programs. Mapped institutional policies into system rules.

Relevance to AI
Understood first-hand how alert fatigue, override rates, and institutional trust interact when automated recommendations meet clinical judgment.
Constraint
Configuration only — did not build the decision support models themselves.
Dr. Smith's Office

Doximity Dialer — Product Lead

CommunicationRegulatedScaleHIPAA

Early product lead for Doximity Dialer — a regulated clinician communication product. 110M+ calls from 300K+ active clinicians. Epic Haiku integration, HIPAA compliance.

Relevance to AI
Shipped a product where compliance, reliability, and user trust were non-negotiable. Managed the gap between what the product promised and what institutions would adopt.
Outcome
110M+ calls, 4.8-star iOS rating (up from 3.7).

Transcarent — Director of Product

Care navigationUtilizationOperations

Led product across 4 specialty-care programs (Surgery, Urgent Care, Behavioral Health, Oncology). Designed utilization rules that reduced avoidable escalations while keeping high-risk cases clinician-led.

How this informs AI work
This experience informs my approach to rule-based routing, human review, and fallback design for AI-assisted workflows.
Outcome
Completed care plans as success metric (not engagement). Unified cross-program member record for consistent triage.

CancerCompass / CTCA — Director of Digital Products

OncologyContentContent moderation

Led product strategy and $2M execution for an oncology platform serving 30MM annual visitors. Built feedback loops with nursing teams to tune triage rules and reduce escalations.

How this informs AI work
Clinical content review and nursing feedback loops inform my approach to review gates for AI-assisted content systems.
Outcome
25% bounce rate reduction, 267% chat conversion increase. Platform later acquired by City of Hope.

Andwise — Co-Founder and CEO

FounderRegulatedRecommendation

Shipped regulated recommendation and review workflows for physician financial decisions. Combined legal oversight (contract analysis) with community-vetted provider directories.

Analogous AI deployment problem
AI-assisted guidance similarly needs explicit scope, qualified review, and an accountable escalation path.
Outcome
$240K raised, 1,200+ physician users, 50+ medical advisory board.