Sourced to artifacts, not self-assessment

Anyone can describe how they work. I had it read off the artifacts.

Not a stranger model with a clever prompt. The one I build with every day, with standing access to every repository, every incident write-up, every operating rule, every handoff, and its own ledger of 8,312 Claude Code sessions since March 10, 2026. The first version of this page was read off 681 ChatGPT working sessions exported April 8, 2026. By September that ledger held 10 times as many messages dated after the export as before it, so I asked the same question again: what does the record actually show about how I operate. What follows is the September answer, edited only to remove private details. None of it is self-assessment.

The finding it kept returning to

He does not trust his memory, his attention, a green check, or an agent's word, and he builds the thing that would catch each one lying.

Not a trait he claims. A pattern visible in 26 dated postmortems, the rules they produced, and the hooks that refuse a command before it runs, most of them written the day something went wrong.

The Inputs

What it was reading.

Not a questionnaire and not a personality quiz. The inputs were the working record itself, most of it written down at the time for reasons that had nothing to do with this page. Every count below was measured on September 20, 2026, and each one is a floor.

Every repository behind the products, commit history included: 87 of them once clones are deduplicated, carrying 64 per-repository instruction files that an agent loads before it touches anything.
The incident log. 26 dated postmortems, 21 of them written since the April export, each one naming the rule it produced. The file's own rule is that no entry is ever deleted.
The rulebook. The standing operating rules that govern how the work gets done, nearly all of them written the day something went wrong. The global file is held at a fixed size by a lint that fails when it grows, so anything added has to be paid for with a removal. 7 more rule files load themselves when a matching file is opened, and a frozen archive of 52 holds the rules as they stood before August 22.
The enforcement. 6 hooks wired into the tooling itself. They run when a command, a browser action or a subagent starts: some inject the rule that applies, and the guards refuse what the rules forbid before it runs. Beside them the lints: copy, secrets, robots, row-level security, and the size of the rulebook itself.
229 handoff documents, one session writing to the next, each opening with a status line so a stale one says so before it misdirects.
The Claude Code record: 8,312 sessions on 2 machines since March 10, 2026, with Claude Code's own ledger of messages, tool calls and tokens. Counted below, with its counting rules.
And the origin. 681 ChatGPT working sessions, 14,077 messages, exported April 8, 2026, which the first version of this page was read off. Kept as the baseline the growth is measured against, not re-read.

Since The Export

The Claude Code record.

This is the corpus the September read came from, and none of it existed when the April one was written. Claude Code's ledger for the laptop holds 149,647 messages dated before the export, across 23 logged days, and 1,493,528 since, across 139: 10 times the volume. The windows are not the same length, so the number that matters is the one that does not care: each prompt he typed moved 22.4 messages of work before the export and moves 70.8 now. The difference is agents running his rules without him, and that is itself part of the record.

  • 8,312Claude Code sessions, 2 machines
  • 2,018sessions he drove by hand
  • 6,294agent sessions he scheduled or dispatched
  • 10,720subagent transcripts
  • 27,760prompts typed
  • 1,643,175messages in the transcripts
  • 664,889tool calls
  • 347Btokens through the model, cache reads included
  • 270sessions on the busiest day, September 6, 2026
The ledger by month, through September 19, 2026
MonthSessionsPrompts typedMessagesTool calls
Mar 20261935,03589,76192,790
Apr 20266846,266266,861134,061
May 20261933,567130,19463,473
Jun 20261012,05729,83410,142
Jul 20261,0152,134182,65657,306
Aug 20262,8075,659532,936159,067
Sep 20261,6613,042410,933148,050

Counting rules. A session he drove is one with at least one prompt typed into the prompt box, taken from the log Claude Code keeps on each machine since March 10, 2026. Agent sessions are everything else Claude Code opened on his machines: scheduled runs, hooks, workflows, subagent swarms. Messages, tool calls and tokens are Claude Code's own ledger for the laptop, through September 19, 2026, and the table above is that ledger by month, so its sessions are the ledger's count and its last row stops there. The ledger began six days after the prompt log, which is why its March row reads short. The Mac Mini's figures are from September 20, 2026. claude.ai chat and cloud sessions are not counted at all. Every figure is a floor: the tooling prunes its own history, so each number is the smallest defensible reading of the record as of September 20, 2026.

The Record

What the record shows.

Nine behaviors it flagged as consistent across the record since the export, each one traced back to something that exists: an incident entry, a rule, a hook, a script. Under each, why it matters to the work itself.

01

Runs a fleet, not a chat.

In April the record was one person typing into one assistant. In September, 6,294 of 8,312 Claude Code sessions are ones he never typed into: scheduled runs, hooks, workflows, subagent swarms dispatched with a brief and a success condition. On the busiest day the ledger shows 270 sessions. Each prompt he types now moves 3.2 times the work it did before the export, and the gap between those two numbers is the operating model: he stopped doing the work by hand and started writing the rules the agents do it under.

Why it matters: the ceiling on a solo builder was always hours in the day. He moved it to how well the rules are written.

02

Writes the postmortem the day it happens, then makes it enforce itself.

The incident log has 26 entries and 21 of them are dated after the export. Each one names the rule it produced, and the rules do not stay prose. A deploy from a stale tree is refused by a hook before it runs. A recursive search over a folder that costs him a password dialog per subfolder was denied by a hook the same day it happened, and that hook then denied the incident write-up itself, minutes later, because the entry quoted the offending command. Nothing in the log is ever deleted, including the entries that later turned out to be wrong. Those carry their corrections inline.

Why it matters: the same mistake is not available to make twice, and the record of having made it once stays legible.

03

Assumes the instrument is lying.

The most repeated rule since April is that a check which returns nothing is broken until proven otherwise. It was written after one session believed seven empty results in a row, one of which wrote an empty secret to a live endpoint. Its corollary arrived in September: a test or a guard is not evidence until it has been watched failing. Measured that week, nine of his own tests passed against deliberately broken code, and a backup alarm built in shell would have watched 24 silent nights and said nothing, because the trap it relied on does not fire for the exact failure it was written to catch.

Why it matters: a false negative looks like good news, and nobody investigates good news. He does.

04

Monitors the outcome, not the machinery.

Reach on his own Threads account fell 90% in July while every monitor read green, because every monitor graded his machinery and none graded the platform's response. A backup wrote ok once and stayed green through 24 days of failure. The Mac Mini went dark for 11 hours with two critical alerts delivered and buried under 27 routine pushes. Each one produced the same question, now asked of every alarm before it is trusted: if the thing this watches broke right now, what in the data it reads would change. If the answer is nothing, the alarm is decoration.

Why it matters: a green dashboard is the most expensive thing to own when it is wrong, because it ends the search.

05

Gates on audience, not on action.

The rule that protects production used to be worded as a prohibition, and agents pattern-matched on the headline: touched code resolved to needs permission, and a proven bug fix shipped as an unverified claim because the session would not build. He rewrote the rule around one question, does the result reach anyone besides him. Builds, tests, local runs, previews and commits are free and listed as never-ask. Anything that reaches a stranger waits for an explicit ship it. In the record a safety rule is a piece of interface design, and where the emphasis sits is the behavior you get.

Why it matters: over-caution is invisible. It looks like work that was never needed. He treats it as a bug with a cost.

06

Refuses the word all.

The domain page said the whole portfolio and it was not: the registrar held more, and the count moved the same week. Now every surface says the count, derived from the inventory so it cannot be wrong, and whole, every and full are banned from it. The same reflex runs through the rest of the record. Every figure on this page is labeled a floor. A machine the census cannot reach keeps its last snapshot and its own date rather than silently shrinking the total. A sweep is reported as found, changed, and the gap by name, because updated the call sites with no count is how nine of twelve ships.

Why it matters: a number with its counting rule attached survives contact with someone who checks.

07

Pays for what he builds, and prices from the invoice.

A wrong price constant, $0.01 recorded against a real $0.14, hid a $172 image loop for six days because the in-app tracker logged under $1 a day while the vendor billed $13. Since then every model is priced from a settled charge, never a pricing page, a fallback may never cost materially more than the primary, no unpriced model goes in a chain, and the hard cap sits at the provider, the one control a bug in his own code cannot defeat. The tradeoff between cost, latency and quality is still not a slide he has seen. It is a bill he has paid, and now one he can read.

Why it matters: usage-based economics is a thing most people have read about. He has been on the paying end across every product he owns.

08

Keeps one source of truth and deletes the copies.

Every Claude Code figure on this site comes from one file written by one script, and a test bans the retired numbers by name. The share cards are rendered from the same file as the page titles, so a card cannot disagree with its own tag. The rulebook is held at a fixed size by a lint. When a figure on this page was found five times too low in September, the fix was not a new number: it was a counting rule printed beside it and a script that fails the moment the page overstates the archive.

Why it matters: anything that can drift, will. He makes drift impossible instead of watching for it.

09

Ships to production, and verifies on the stranger's surface.

The portfolio is still not a folder of prototypes: products live at their own domains, apps on the App Store, real people on the other end. What changed since April is where the verification stands. Verified now means the deployed URL not the source, the shipped bundle not the commit, the public record not the confirmation page, and it names its auth state and viewport or it is not a verdict. A sixteen-item wave once shipped as verified when every check in it had been signed out. That is the incident the rule came from.

Why it matters: anyone can run a demo. Knowing what breaks between the demo and production is what makes a promise safe to make.

The Limits

What it would not say.

This is only worth publishing because of where it stops.

Artifacts only. It discounted marketing copy, including the copy on this site. If a claim existed as a sentence and nowhere else, it did not count.
Architecture over volume. Nearly all of the raw output since the export is agent-produced, so it credited the rules, the gates and the judgment calls inside them rather than the line count. Deciding what to build and what it has to survive is the harder half anyway.
Solo evidence. Nearly the entire record is Jack working alone or directing agents. It has very little to say about how he behaves inside a team, and it did not pretend otherwise. That gap is real and it is left as one.
It re-read, it did not re-grade. Where the April behaviors still hold they are not repeated for the sake of a longer list. Decides in the room, codeswitches on purpose, adversarial with his own claims: all three are still in the record. The nine above are what the record since the export added or sharpened.
One record, not a benchmark. Every finding here is a description of behavior with a source attached. None of it is a grade, and there is nothing in it to compare against anyone else.
Nobody is a reliable narrator about themselves. That is why this page is built out of artifacts instead of adjectives.
Why this page exists

First generated in 2026 from the archive as it stood that April. Re-read on September 20, 2026 off the record since, which by then ran to 10 times the messages that came before it. Private details removed, nothing else rewritten. The closing note from the April analysis still stands, and the record since is mostly evidence for it: write down what broke, turn it into a rule, then make the rule enforce itself.