Applied AI systems for consequential work

AI that survives contact with the real world.

I design and operate agentic systems for real workflows. Tools run under explicit permissions, state survives a failure, and a person with authority sits at every consequential step. Each production miss becomes a test the next release has to pass.

Substrate

What the systems are made of.

Agent runtime

Every tool Stem can run carries a declared schema. Anything that writes, runs a shell, or destroys goes through the permission gate: destructive tools start with an empty grant, and an always-allow lasts one session and dies with it.

The permission gate

Evaluation

A failure that matters becomes a case the next version has to pass, and the suite only tightens. The judge gets judged too: a stale gold file had a model benched for three weeks at 41%; re-reading the transcript against frozen fixtures put it at 91%.

Smarter every time it fails

Production intelligence

19,124 production events run through the redaction gate with zero confirmed leaks. One trace audit found 1.5 million tokens in a month spent re-reading a single file; a one-line rule in the operating doc closed it the same day.

Zero confirmed leaks across 19,124 events

Workflow systems

The Review Desk runs every new bid against ten years of the organization's documented estimating mistakes. A finding without source evidence never reaches the estimator, and commercial scope stays with the estimator. $221,915.70 surfaced across $3.44M; zero validated findings escaped into a sent quote.

How the system divides responsibility
01 / Operating theses

Principles earned in production.

  1. Thesis 01

    Reconciliation before reasoning.

    In many AI workflows the hardest part is establishing the true state of the work when several sources disagree. The review desk exists because quotes, takeoffs, pricing, drawings, and service rules could each look correct on their own while the final bid was still wrong.

    See the system
  2. Thesis 02

    Failure should become executable knowledge.

    Stem turns observed failures into requirements, regression cases, routing changes, and safeguards. What began as individual mistakes has become a permanent evaluation suite that future changes must pass.

    Read the field note
  3. Thesis 03

    Autonomy should follow consequence.

    Whether an AI system can act matters less than whether the cost of being wrong justifies letting it act without evidence, deterministic validation, or human approval.

    Read the field note
02 / Field notes

What happened in production, and what changed because of it.

All field notes

03 / Evidence from the field

Three systems in production.

01 / Commercial estimating

Ten years of the company's bid errors, turned into a detector that runs on every new bid

Reconciliation · evidence provenance · human-in-the-loop AI

I pulled every estimating mistake the organization had documented since 2015, identified the failure patterns that repeatedly cost margin, and built a system that checks every new bid for them before it ships. Every finding is tied to source evidence; the estimator makes the final call.

$221,915.70 surfaced · $3.44M reviewed · 6.4% verified value catch rate · zero validated findings escaped into a sent quote

Read the case

02 / Local-first AI

Stem

Agent runtime · evaluation infrastructure · local-first AI

A production AI system built around a product constraint: sensitive personal and professional work should remain private without giving up access to capable models and modern agent workflows. I own the product end to end, including requirements, interaction design, model routing, orchestration, evaluation, safeguards, and production iteration.

19,124 production events · zero confirmed leaks · 96/96 self-test gate

See the product

03 / Workflow systems

The quote import tool

Deterministic workflow automation · state transformation · production desktop software

A roughly 50-person organization was moving the same information manually through three disconnected systems, creating duplicate work and opportunities for error. I took the product from workflow discovery through prototype, beta, production rollout, and post-launch expansion, replacing both manual transfer points with a deterministic browse-review-save workflow designed around a locked-down enterprise environment.

30 to 60 hrs/wk returned · 12 releases · no rollbacks

Read the case

Thirteen working products across iOS and the web sit behind these systems: creative tools, audio instruments, utilities, games, research software, and internal tools. I build directly because a working version reveals what a requirements document cannot. See the fleet.

Built in the field.

Ten years leading commercial developments where a missed milestone carried liquidated damages. Then I started turning those operating problems into software.

Over the last four years I have owned AI products from discovery and 0-to-1 development through launch, evaluation, and post-launch iteration. How I work and the full record are on the About page.

FutureKind is the independent product practice where much of this work is built, tested, and deployed.