A mental-health journalling platform where the AI is not allowed to make things up. When an entry’s sentiment drops, the feedback has to quote clinical sources — WHO, APA, NHS, NIMH — rather than generate advice freely. Alongside it, a Random Forest reads the last fourteen days of behaviour and forecasts the next three, so a difficult stretch can be flagged before it arrives. Every layer of it — machine learning, backend, frontend, mobile, infrastructure — was built by one person.
Try it, and read the code
Live at heartbox.tw. No registration — sign in with a test account.
Three accounts show how the system behaves under different emotional patterns. These are the English-language accounts; the interface follows the account, so signing in with one of these gives you the site in English:
| Account | Password | Pattern | Worth looking at |
|---|---|---|---|
test1_en | test1_en | Volatile | How the charts render swings, and how the physiological correlations surface sleep and activity effects |
test2_en | test2_en | Positive | How progress is shown when the trend is good |
test3_en | test3_en | Negative | What the system does when mood stays low — the support flow and the knowledge-base citations |
Source: github.com/alanlin0604/HeartBox
Problem
Roughly 5.7% of adults worldwide live with depression — about 332 million people (WHO) — and the rate of psychological distress among young people in Taiwan has been climbing. Asking for help is itself the obstacle: counselling capacity is limited, waits are long, and stigma stops many people taking the first step at all.
Mood-tracking apps do not fill that gap, and not simply because they lack features:
- The feedback is shallow. They record a mood score and stop there.
- The AI invents things. A general-purpose LLM will produce psychological advice with no basis behind it. In this domain that is not a harmless error.
- Sensitive text is stored in the clear. A journal is among the most private writing a person produces.
- Everything looks backwards. You can see the low patch you already went through, not the one arriving.
HeartBox addresses the space before professional help becomes necessary: a self-awareness tool that is grounded, private, and able to show what is coming. It does not replace counselling; it lowers the cost of noticing that something is wrong.
Constraints
These shaped every technical decision that follows:
- There is almost no data per user. Hundreds of entries, not tens of thousands. Anything that needs volume to train was out from the start.
- The data is unusually sensitive. Sending journal text to a third-party API is a hard decision to justify.
- Advice cannot be invented. When a chatbot is wrong the cost is embarrassment. Here it is not.
- Feedback has to be immediate. The analysis appears when the entry is saved, not later.
- Individual variation defeats cross-user comparison. One person’s stress is driven by poor sleep, another’s by inactivity. An “average user” means nothing here.
- One person builds and runs it. Anything needing constant tuning or babysitting becomes a liability rather than a feature.
Key decisions
This is the part of the page that matters. Each decision is written as situation → alternatives → trade-off → outcome.
Why retrieval-augmented generation instead of prompting an LLM directly?
- Situation
- After someone writes a low entry, the system has to respond. A general-purpose LLM will produce advice that reads as authoritative and may be entirely invented — and because the tone is warm and the structure is confident, it is more likely to be believed, not less.
- Alternatives considered
- Call an LLM directly and ask it in the prompt to be careful
- Drop generation entirely and return fixed, pre-written responses
- Retrieval-augmented generation: find relevant clinical passages first, then require the model to answer from them
- Trade-off
- RAG adds a vector lookup and a longer prompt to every response, so it is slower and more code to get wrong. It also makes the knowledge base a hard ceiling on answer quality. Canned responses are perfectly safe but generic enough that people stop reading them after the second one.
- Decision
- Retrieval-augmented generation. The knowledge base holds seven documents from international bodies — WHO's stress-management guidance, its mental-health action plan and Doing What Matters in Times of Stress; APA on stress-management technique and building resilience; NHS guidance on mental wellbeing; NIMH on coping with stress — indexed with BGE-M3 embeddings, which turn each passage into a string of numbers standing for its meaning so it can be searched by similarity rather than by keyword. Retrieval returns the three closest passages and the model must stay inside them. The cost is latency and complexity; what it buys is that any given sentence of advice can be traced to where it came from.
The pipeline is five steps:
- The user writes an entry
- The model scores sentiment on a −1 to +1 scale
- Below −0.4, and only below −0.4, retrieval runs
- ChromaDB returns the three most relevant passages from the seven documents
- The model composes feedback from what was retrieved
Why is the retrieval threshold at −0.4?
- Situation
- Not every entry warrants citing clinical literature. 'Tiring day' and 'I do not want to do anything anymore' both score as negative, but they call for very different responses. Where the threshold sits decides whether the system feels attentive or preachy.
- Alternatives considered
- No threshold — run retrieval on every entry
- A high threshold such as −0.2, erring towards triggering
- A low threshold such as −0.6, intervening only when things are clearly bad
- Trade-off
- Set it too high and the system answers everyday grumbling with psychological literature, which reads as lecturing and gets tuned out fast — the mirror image of a missed warning, and just as damaging to trust. Set it too low and the quietly-worded entries, the ones most worth catching, slip past. Running it on everything pays latency and GPU cost on entries that never needed it.
- Decision
- −0.4. It lets clearly negative entries through while leaving ordinary complaints alone. Worth being honest about: the number came from judgement against real entries, not from a sweep. No sensitivity analysis backs it, and that is the first experiment I would add.
Why Random Forest rather than XGBoost or an LSTM?
- Situation
- Forecasting mood from a behavioural time series points straight at an LSTM. The actual conditions point elsewhere: a few hundred rows per user, a prediction that has to return while the entry is still being saved, and users who will reasonably ask why the system thinks their stress is about to rise.
- Alternatives considered
- Linear regression — simplest, most explainable
- A single decision tree — completely transparent rules
- Random Forest — an ensemble of 100 trees
- XGBoost — usually the most accurate once tuned
- An LSTM — the intuitive choice for sequential data
- Trade-off
- Passing on an LSTM gives up whatever long-range temporal structure it might have found, and the accuracy ceiling it would reach once the dataset grows. Passing on XGBoost gives up the few points tuning usually buys. In exchange: it trains on a few hundred rows, predicts in under 50 ms on CPU, explains itself through feature importance — the model's own account of which inputs mattered most — and needs essentially no tuning.
- Decision
- Random Forest, 100 trees. Regression targets — sentiment and stress — average across the trees; the high-stress classifier averages the probability each tree assigns. It is the only candidate that satisfied all four constraints that actually bound this project at once.
| Algorithm | Data needed | Interpretability | Inference cost | Overfitting risk | Accuracy and stability |
|---|---|---|---|---|---|
| Linear regression | Very little | High (coefficients) | Negligible | Tends to underfit | Capped by the linearity assumption |
| Single decision tree | Very little | High (explicit rules) | Negligible | High | Unstable |
| Random Forest★ chosen | A few hundred rows | Medium-high (feature importance) | Low — sub-50 ms, CPU only | Low (bagging resists it) | High, stable, near zero tuning |
| XGBoost | Moderate | Medium-high (feature importance) | Low | Medium (needs tuning) | High once tuned, often best |
| LSTM | Tens of thousands of rows | Low (opaque) | Medium (needs a DL runtime) | High on small data | High with enough data |
Why tune for 88% recall rather than precision?
- Situation
- High-stress-day warning is a classification problem. Recall is the share of genuine high-stress days the system catches; precision is the share of its warnings that turn out to be right. The two pull against each other — warn more freely and you catch more, but you are also wrong more often. Where the threshold sits comes down to which error costs more.
- Alternatives considered
- Balance the two and maximise their harmonic mean (F1)
- Favour precision — warn only when confident, and avoid nagging
- Favour recall — accept false alarms rather than miss a real one
- Trade-off
- Favouring recall means more false alarms: a warning that a hard few days may be coming, followed by nothing happening. Repeated often enough that devalues every warning, which is a real cost and not one to wave away.
- Decision
- Recall, landing at 88%. The two errors are not symmetric: a false alarm costs one unnecessary nudge, a missed one costs a low patch nobody caught. In this domain that asymmetry is large enough to decide the threshold on its own. The wording follows the same logic — the system says something worth keeping an eye on rather than issuing a warning, so a false positive lands gently.
Why self-host inference instead of calling a commercial API?
- Situation
- Calling a hosted API is by far the easier path: no GPU to run, no model to deploy, and generally better output. But the text being processed is a person's private journal.
- Alternatives considered
- Call a commercial LLM API — fastest to build, least to operate
- Call a commercial API after de-identifying the text
- Self-host open-weight models so the data never leaves a controlled environment
- Trade-off
- Self-hosting is expensive in the ways that matter: a GPU to provide, model serving to build, and outages to own. Open-weight 7B models also generate noticeably weaker Chinese than a large commercial model. De-identification sounds like the middle path, but a journal is identifying by nature — the details of someone's life do not anonymise.
- Decision
- Self-hosted. Open-weight Qwen2.5-7B-Instruct runs on my own GPU behind FastAPI, reached through a Cloudflare Tunnel. The generation-quality gap against a commercial API is a real price. But 'journal text has never left an environment I control' is a property this product does not get to trade away.
Why split consent into three separate steps?
- Situation
- The system collects journal text, the sentiment and stress scores derived from it, and optionally health and sleep data. GDPR defines consent as specific to a purpose (Art. 4(11)), expects distinct processing operations to be separately consentable (Recital 43), and in Art. 7(4) weighs against any service made conditional on consent the contract does not require. Taiwan's Personal Data Protection Act, Art. 7(2), likewise requires consent to out-of-purpose use to be given on its own. A single 'I agree to everything' checkbox does not meet it.
- Alternatives considered
- One consent form covering everything
- Two stages: terms of service, then data use
- Three separate stages: data use, AI-training consent, age confirmation
- Trade-off
- Three steps instead of one click will cost sign-up conversion, and letting people refuse AI-training use means deliberately giving up training data that would otherwise be available. Both losses are real.
- Decision
- Three separate stages. The load-bearing detail is that AI-training consent is its own checkbox and refusing it leaves every feature available — if refusal degraded the product, the consent would not be freely given in the legal sense. Users aged 13 to 17 go through parental verification: the system emails a confirmation link, and AI features unlock only once it is followed. The conversion cost is what this decision is worth paying.
Why compare users against themselves rather than against a baseline?
- Situation
- Once the forecast worked, the next question was how to show someone whether they are improving. The obvious answer is to compare against other users — but variation between people is so wide that the comparison carries almost no meaning. A naturally lower baseline mood is not a warning sign.
- Alternatives considered
- Cross-user percentile: 'your score is higher than 70% of users'
- A fixed absolute threshold: 'below −0.3 counts as bad'
- A personal baseline: the user's own first week
- Trade-off
- A personal baseline cannot answer 'how do I compare to other people', and it has nothing to say to a user who just signed up — with too little history the honest output is no output, which is a visible gap in the product experience.
- Decision
- Personal baseline. Progress is measured against the user's own first week, and an upward arrow always means improvement so nothing has to be memorised. Below fourteen days of history the interface shows how many days remain before the comparison unlocks, rather than computing a number that would not mean anything. This is the decision I am happiest with in the whole project — it admits the data is thin instead of papering over it with something that looks rigorous.
Results





Feature engineering produced 53 features: twelve behavioural and physiological metrics — mean sentiment, mean and peak stress, entry count, sleep duration and quality, deep-sleep proportion, step count, active minutes, HRV, resting heart rate, and bedtime variance — each expanded across 1, 3, 7 and 14-day windows for 48 lag features, plus 5 calendar features such as day of week and journalling streak.
What these numbers show, and what they do not. They are offline cross-validation results: evidence that the model learns something predictive from behavioural history. They are not evidence of clinical validity. There was no controlled trial, and nothing here demonstrates that the warnings improved anyone’s mental health. That distinction matters more than the figures do, and good-looking numbers should not be allowed to blur it.
What I’d do differently
-
Validate on a temporal split, not random 5-fold. The feature set is dominated by lag variables, so a random split lets information from the same stretch of time land in both training and validation, and the metrics are probably optimistic because of it. Training on earlier periods and predicting later ones is the honest evaluation. This is the project’s biggest methodological weakness and the first thing I would fix.
-
The −0.4 threshold was never swept. It is a defensible value chosen against real entries, but I never quantified how the false-trigger rate moves at −0.3 or −0.5. That experiment was affordable; I spent the time on features instead.
-
Retrieval quality itself was never measured. The entire value of the RAG pipeline rests on the retrieved passages being relevant, yet I built no retrieval evaluation set and never measured top-3 hit rate. A model citing a source is not the same as a model citing the right source.
-
Too many features, diluting the core. A community feed, friends, courses, achievements, breathing exercises — all complete, none of them connected to the two things that actually carry the project. Given the time again I would put it into retrieval evaluation and prospective validation rather than into the hundred-and-third achievement badge.
-
I underestimated the operational cost of self-hosting. I still think the decision was right, but when the GPU service goes down every AI feature goes with it, and I built no degraded mode. At minimum it should fall back to plain journalling without analysis rather than showing an error.
Stack
| Layer | Technology |
|---|---|
| Frontend | React ・ Vite ・ Tailwind CSS ・ Recharts |
| Backend | Django REST Framework ・ Django Channels (WebSocket) |
| AI | Qwen2.5-7B-Instruct ・ BGE-M3 ・ LangChain ・ ChromaDB |
| Machine learning | scikit-learn (Random Forest) |
| Database | PostgreSQL, journal fields Fernet-encrypted (AES-128-CBC + HMAC-SHA256) |
| Mobile | Capacitor (Android) |
| Deployment | Cloudflare Pages ・ Cloudflare Tunnel ・ Google Cloud Run |
On security and privacy: journal contents are stored under Fernet symmetric encryption with keys injected from the environment and held separately from the database, so a database leak on its own does not recover them. Sign-in supports two-factor authentication and the site is HTTPS throughout. Crisis-keyword detection runs across journalling, AI chat and the anonymous community, and surfaces national helplines immediately when it fires.