Sindri
Open the app →

Case study / building sindri

A second brain, built by arguing with the thing helping build it.

Sindri captures, remembers, and resurfaces — a personal memory system built because forgetting things has a real cost. This page is not a highlight reel. It's the running record of where the design was wrong first: the assumptions that turned out false, the shortcuts that turned out to be risks, and the moments a plan changed because someone — human or model — checked instead of assumed. Six entries logged so far, one at a time, by hand.

Human — judgment, priorities, lived experience with the toolAI — what got built, measured, or caught on its own

Every entry below is written by hand and reviewed like any other change to this codebase — nothing here is generated from commit messages.

2026 · 08 · 23
measured, not assumed

Deleting a planned feature because nobody had actually tried it

The original design for reading a research goal's watch list planned a dedicated tool for reaching X, reasoned from one premise: X is a login wall to any real extractor, so it could only ever be discovery-only, by design, forever. That premise was never tested against the extractor actually in use — it was a guess about X's reputation, not a measurement. One probe against the "proper" API found half of it dead and its replacement absurdly priced — roughly $0.85 a call, 28× the token cost, for a narrative no other tool could plug into. A second probe tried the boring alternative already sitting in the codebase: an ordinary scoped web search plus the page-fetch tool. It reached real posts, quote-verified, 19 times out of 19. The planned tool was deleted rather than shipped behind a flag.

HumanPushed back on "a login wall, forever" as a claim worth testing before a feature got built around it.
AIWrote both throwaway probes, measured both paths for real, and reversed the design based on what came back.
2026 · 08 · 19
honesty

The number that was worth publishing precisely because it looks bad

A retro feature was built to say whether a research goal is actually working, using the keep/reject verdicts a user has given its findings over time. Checked against the real corpus — 27 findings — exactly zero had ever been rejected as "off-target" or "inaccurate." An 8-for-8 keep rate on any single goal reads as proof it's well-tuned, but it just as easily means nobody has judged it harshly yet. That caveat now ships with every retro number the page renders, computed fresh from the corpus each time rather than written once as static disclaimer text that could quietly go stale.

HumanSet the bar: a summary can't just report counts, it has to say what the counts can't prove.
AIQueried the real corpus for the actual number instead of estimating it, and wired the caveat to live data.
c. 2026 · 08
self-correcting

A hedge that turned out to be right twice over

The morning digest was told: if a number looks wrong to you, send it anyway and say what you noticed underneath. On its first real delivery, the model used that instruction — it flagged that it couldn't verify two items had "arrived overnight." It was right on both counts. The 24-hour window at a 7am send actually reached back through the whole previous day, so the label was false; and "does this look wrong to you" had no way for the model to actually check, so the honest answer was always going to be a hedge. Both got fixed: the heading now says "since yesterday" with real timestamps, and the prompt now points at a tool call instead of asking for a feeling.

HumanWrote the original instruction, and treated the model's hedge as a signal worth investigating rather than noise.
AISurfaced the uncertainty instead of asserting a clean number, which is what exposed the actual bug underneath it.
2026 · 08 · 01
risk call

Two cards for one obligation — link, don't merge

Gmail and Slack each run their own sweep for things the user still owes someone, and neither can see the other's work. Twice in production, both raised a card for the exact same real ask — once with no shared word, link, or name between the two. Merging them into one card was the obvious fix, and it was rejected: a merged card needs a closing rule nobody could write, since each sweep can only observe its own channel going quiet. Worse, a wrong merge silently deletes one of two real obligations, while a duplicate left standing only costs a glance. The asymmetry decided it before any detector existed.

HumanWeighed which failure was worse — a visible duplicate or a silent loss — and picked the direction that fails safe.
AIBuilt the link as symmetric by construction, then wrote a test that fails if a future change lets the two sides drift.
2026 · 07 · 31
premise

The assumption nobody had actually checked

Two separate planning documents had argued against building a web browsing surface for notes, on the same reasoning: "Obsidian is already the browse surface." Neither argument was ever true in practice — the user had tried Obsidian repeatedly and didn't like it. Said plainly, that one fact flipped the browsing surface from "unnecessary" to "the actual next feature." It shipped days later.

HumanReplaced an architectural assumption with the one fact that actually mattered: real usage.
AIRetired the old argument in the documentation itself, so the next design decision wouldn't inherit it again.
2026 · 07 · 29
self-audit

Turning the project's own rule on its own code

One rule this project had already learned the hard way: a name describing what code is supposed to do is not evidence that it does it. Turned back on the project's own output in a single session, it found five identifiers that promised more than they delivered — a field marked "optional" that three queries actually required, a delete method with no caller outside its own test, an inbox heading claiming "untriaged" while quietly ignoring status, an auth field read off a type that didn't have it, and a safeguard called "the fabrication gate" in four places despite doing something narrower than that name implies.

HumanSet the standard the audit ran against, forged from three earlier integration failures with outside vendors.
AIRead its own recent code the way it would read a stranger's undocumented API, instead of trusting its own names.

This log only runs one direction — forward. Nothing here gets deleted when a later entry complicates it, because the earlier wrong turn is the part worth keeping. c. marks a date the source material didn't record exactly; everything else is real.