Case study / 01
Nura
A guided space for the night after a fight.
The moment
The moment it's built for
It's 11pm. Someone has just had the same argument with their partner for the fourth time this year. Their friends will take a side. A chatbot will hand them an empty box and wait for them to explain themselves, at the exact moment they can least do it. Therapy is weeks away and treats this as something clinical.
Nura is for that person, at that hour. They open the app, tap Something happened, and Nura asks one question at a time for about ten minutes. It asks how they are before it asks about the fight. It notices the word they keep using. When they're done, a card deals onto the screen: what happened, what they felt, and on the back, one observation in their own words and one small thing to try tomorrow.
The next evening, a single quiet notification asks how it went. After a few sessions, a screen called the Mirror shows them the loop they keep ending up in, with a quote from them under every step.
The bet
Why it can win
Anyone can wrap a model in a chat window. Nura's edge is what it remembers and how carefully it earns the right to say something. Every session becomes a structured record. The records become evidence. By the tenth session Nura knows things about this person's pattern that a fresh chatbot never will, and the person can see exactly which of their own words each claim rests on.
The business model follows the same values. There are three free sessions, then one plan. The paywall appears only at the moment someone keeps their third card, never in the middle of a conversation. A session that ended with crisis resources on screen never counts toward the limit. People keep read access to everything they wrote, forever, whether they pay or not.
The first users are 25 to 34, in relationships of a year or more, starting with testers in Kathmandu and remote teams. The product treats family involvement in relationship decisions as normal context, which matters in South Asia and is something most Western tools get wrong.
Decisions
Five decisions that shaped it
01 / five
The model reads, the code decides
Language models in Nura do four jobs: write the next line, pull structure out of a finished session, screen messages for safety, and grade test runs. Everything that has consequences for a person is ordinary, tested code. Which question comes next, when a pattern is allowed on screen, when someone reaches the paywall, when a conversation stops for safety.
That split is why one of the product's hardest rules can be guaranteed. Someone whose problem is a real difference in values (kids, where to live, religion) must never be handed communication exercises, as if better listening would fix it. That help is gated behind evidence, and a unit test runs a full values-conflict history through every session type to prove the exercises can't be reached.
02 / five
No quote, no claim
The worst thing Nura could do is tell someone "you said this" about words they never said. So when a session ends, every quote the model pulls out is matched against the saved transcript before it can reach a card or the Mirror. A quote that doesn't match is dropped in code.
The end-to-end test includes one evening where the model invents a quote on purpose. Switch the check off and that invented sentence lands on a card, and two tests fail. Patterns are recalculated from the full history every time, so when someone deletes a session, everything Nura believed because of it disappears too.
03 / five
Every message is screened before anything is shown
Some people will write about self-harm, abuse or a partner who controls their money. Those moments need the conversation to stop and real help to appear, calmly.
The first design scanned for alarming keywords. I wrote 69 test messages the way people actually say these things, and the scan caught 21 of them. Now every message is scored by Jev, a dedicated safety classifier, while Nura drafts its reply, and nothing reaches the screen until the verdict is in. The thresholds live in code, one set per category. Claude Haiku gives a second opinion where ordinary arguments and real control look alike.
On 145 test messages the design catches 97% of disclosures with 9% false alarms. I wrote up how that went, including what's still unsolved, in the Jev post.
04 / five
Privacy for the person whose partner might pick up the phone
That's the threat model, and it changes a lot of small decisions.
- There's exactly one notification in the whole product, and its data type has no field that could carry anyone's words. It says "A quick check-in when you have a minute" and nothing else.
- Each person's words are encrypted under their own key. Deleting an account destroys that key first, before a single row is touched.
- Error reports and traces pass through a gate that drops anything it doesn't recognize. No message, name or user id ever leaves the server, even inside a crash report.
- Consent is stored with a fingerprint of the exact words the person agreed to. If those words ever change, Nura stops processing and asks again rather than assuming.
05 / five
Measure what helps, refuse what hooks
There are no streaks, no daily reminders and no "you haven't opened Nura in a while". The weekly quality review watches six numbers, such as how long a card takes to arrive and how often people read the back of it. Session length, return rate and streaks are listed in the code as numbers the product must never track, and adding one breaks a test. Sessions that ended in a safety hand-off are kept out of the drop-off rate, so nobody is ever rewarded for making safety quieter.
Under the hood
Under the hood
One repo, two processes
An API and a background worker, with a build rule proving the worker never loads the web server. Seven job queues, retries, dead letters, and sweeps that recover lost jobs without ever sending anything twice.
Memory without vector search
Typed records and a pure function that folds them into competing explanations, each with a confidence score and linked quotes. You can always say why something is in the model’s context.
One gateway for every model call
Through OpenRouter, with zero data retention required on every request. If a provider can’t honor that, the request is refused.
Versioned everything
Prompts, models and product rules are versioned files in the repo. Every card and record stores the version that produced it.
Streaming conversation over SSE
Every generated line is checked against Nura’s voice rules (no diagnosis words, no clichés, no exclamation marks) before it’s shown.
Accounts
15-minute access tokens, rotating refresh tokens, and a full sign-out on every device if a used token ever reappears. Export and deletion of everything, behind a password check.
Billing through RevenueCat
Each webhook is treated as a prompt to re-read the source of truth, so duplicate and late events are harmless.
Process
How I build
I designed Nura's product and architecture and built it with Claude Code as my coding agent. That works only with firm rules, so I wrote them down. Types make impossible states impossible to write. Decisions live in pure functions. Architecture boundaries are enforced by the build. And every rule ships with a test I've watched fail by breaking the code on purpose.
Each phase ends with a written log of what shipped, what changed on the way, and what I chose not to do and why. About 40,000 lines of code and 37,000 lines of tests later, the log is the reason the codebase still makes sense.