Skip to content
Cheng-Han Lin
Role
Solo developer (capstone)
Period
Sep 2025 – Jun 2026
Kind
Capstone project
Scope
ML · backend · frontendmobile · infrastructure
Stack
Django REST FrameworkReactPostgreSQLChromaDBscikit-learn
Links
heartbox.twSource

HeartBox

A mental-health journalling platform built on retrieval-augmented generation, with a Random Forest emotion forecast and self-hosted open-weight model inference. Every layer of it — ML, backend, frontend, mobile, infrastructure — was built by one person.

Results

5-fold cross-validation. 22,796 rows for the regression tasks, 31,720 for high-stress classification.

Sentiment score MAEscore range −1 to +1
0.22
Stress index MAEindex range 0 to 10
1.04
High-stress AUCnear-perfect ranking
0.948
High-stress recalldeliberately biased against misses
88%
The mood-trend card from the HeartBox dashboard — average stress swinging between 0 and 9 in red, average sentiment hovering near 0 in orange, across the past 180 days

A mental-health journalling platform where the AI is not allowed to make things up. When an entry’s sentiment drops, the feedback has to quote clinical sources — WHO, APA, NHS, NIMH — rather than generate advice freely. Alongside it, a Random Forest reads the last fourteen days of behaviour and forecasts the next three, so a difficult stretch can be flagged before it arrives. Every layer of it — machine learning, backend, frontend, mobile, infrastructure — was built by one person.

Try it, and read the code

Live at heartbox.tw. No registration — sign in with a test account.

Three accounts show how the system behaves under different emotional patterns. These are the English-language accounts; the interface follows the account, so signing in with one of these gives you the site in English:

AccountPasswordPatternWorth looking at
test1_entest1_enVolatileHow the charts render swings, and how the physiological correlations surface sleep and activity effects
test2_entest2_enPositiveHow progress is shown when the trend is good
test3_entest3_enNegativeWhat the system does when mood stays low — the support flow and the knowledge-base citations

Source: github.com/alanlin0604/HeartBox

Problem

Roughly 5.7% of adults worldwide live with depression — about 332 million people (WHO) — and the rate of psychological distress among young people in Taiwan has been climbing. Asking for help is itself the obstacle: counselling capacity is limited, waits are long, and stigma stops many people taking the first step at all.

Mood-tracking apps do not fill that gap, and not simply because they lack features:

  • The feedback is shallow. They record a mood score and stop there.
  • The AI invents things. A general-purpose LLM will produce psychological advice with no basis behind it. In this domain that is not a harmless error.
  • Sensitive text is stored in the clear. A journal is among the most private writing a person produces.
  • Everything looks backwards. You can see the low patch you already went through, not the one arriving.

HeartBox addresses the space before professional help becomes necessary: a self-awareness tool that is grounded, private, and able to show what is coming. It does not replace counselling; it lowers the cost of noticing that something is wrong.

Constraints

These shaped every technical decision that follows:

  1. There is almost no data per user. Hundreds of entries, not tens of thousands. Anything that needs volume to train was out from the start.
  2. The data is unusually sensitive. Sending journal text to a third-party API is a hard decision to justify.
  3. Advice cannot be invented. When a chatbot is wrong the cost is embarrassment. Here it is not.
  4. Feedback has to be immediate. The analysis appears when the entry is saved, not later.
  5. Individual variation defeats cross-user comparison. One person’s stress is driven by poor sleep, another’s by inactivity. An “average user” means nothing here.
  6. One person builds and runs it. Anything needing constant tuning or babysitting becomes a liability rather than a feature.

Key decisions

This is the part of the page that matters. Each decision is written as situation → alternatives → trade-off → outcome.

Why retrieval-augmented generation instead of prompting an LLM directly?

Situation
After someone writes a low entry, the system has to respond. A general-purpose LLM will produce advice that reads as authoritative and may be entirely invented — and because the tone is warm and the structure is confident, it is more likely to be believed, not less.
Alternatives considered
  • Call an LLM directly and ask it in the prompt to be careful
  • Drop generation entirely and return fixed, pre-written responses
  • Retrieval-augmented generation: find relevant clinical passages first, then require the model to answer from them
Trade-off
RAG adds a vector lookup and a longer prompt to every response, so it is slower and more code to get wrong. It also makes the knowledge base a hard ceiling on answer quality. Canned responses are perfectly safe but generic enough that people stop reading them after the second one.
Decision
Retrieval-augmented generation. The knowledge base holds seven documents from international bodies — WHO's stress-management guidance, its mental-health action plan and Doing What Matters in Times of Stress; APA on stress-management technique and building resilience; NHS guidance on mental wellbeing; NIMH on coping with stress — indexed with BGE-M3 embeddings, which turn each passage into a string of numbers standing for its meaning so it can be searched by similarity rather than by keyword. Retrieval returns the three closest passages and the model must stay inside them. The cost is latency and complexity; what it buys is that any given sentence of advice can be traced to where it came from.

The pipeline is five steps:

  1. The user writes an entry
  2. The model scores sentiment on a −1 to +1 scale
  3. Below −0.4, and only below −0.4, retrieval runs
  4. ChromaDB returns the three most relevant passages from the seven documents
  5. The model composes feedback from what was retrieved

Why is the retrieval threshold at −0.4?

Situation
Not every entry warrants citing clinical literature. 'Tiring day' and 'I do not want to do anything anymore' both score as negative, but they call for very different responses. Where the threshold sits decides whether the system feels attentive or preachy.
Alternatives considered
  • No threshold — run retrieval on every entry
  • A high threshold such as −0.2, erring towards triggering
  • A low threshold such as −0.6, intervening only when things are clearly bad
Trade-off
Set it too high and the system answers everyday grumbling with psychological literature, which reads as lecturing and gets tuned out fast — the mirror image of a missed warning, and just as damaging to trust. Set it too low and the quietly-worded entries, the ones most worth catching, slip past. Running it on everything pays latency and GPU cost on entries that never needed it.
Decision
−0.4. It lets clearly negative entries through while leaving ordinary complaints alone. Worth being honest about: the number came from judgement against real entries, not from a sweep. No sensitivity analysis backs it, and that is the first experiment I would add.

Why Random Forest rather than XGBoost or an LSTM?

Situation
Forecasting mood from a behavioural time series points straight at an LSTM. The actual conditions point elsewhere: a few hundred rows per user, a prediction that has to return while the entry is still being saved, and users who will reasonably ask why the system thinks their stress is about to rise.
Alternatives considered
  • Linear regression — simplest, most explainable
  • A single decision tree — completely transparent rules
  • Random Forest — an ensemble of 100 trees
  • XGBoost — usually the most accurate once tuned
  • An LSTM — the intuitive choice for sequential data
Trade-off
Passing on an LSTM gives up whatever long-range temporal structure it might have found, and the accuracy ceiling it would reach once the dataset grows. Passing on XGBoost gives up the few points tuning usually buys. In exchange: it trains on a few hundred rows, predicts in under 50 ms on CPU, explains itself through feature importance — the model's own account of which inputs mattered most — and needs essentially no tuning.
Decision
Random Forest, 100 trees. Regression targets — sentiment and stress — average across the trees; the high-stress classifier averages the probability each tree assigns. It is the only candidate that satisfied all four constraints that actually bound this project at once.
Five candidate algorithms compared on data requirements, interpretability, inference cost, overfitting risk and accuracy
AlgorithmData neededInterpretabilityInference costOverfitting riskAccuracy and stability
Linear regressionVery littleHigh (coefficients)NegligibleTends to underfitCapped by the linearity assumption
Single decision treeVery littleHigh (explicit rules)NegligibleHighUnstable
Random Forest★ chosenA few hundred rowsMedium-high (feature importance)Low — sub-50 ms, CPU onlyLow (bagging resists it)High, stable, near zero tuning
XGBoostModerateMedium-high (feature importance)LowMedium (needs tuning)High once tuned, often best
LSTMTens of thousands of rowsLow (opaque)Medium (needs a DL runtime)High on small dataHigh with enough data

Why tune for 88% recall rather than precision?

Situation
High-stress-day warning is a classification problem. Recall is the share of genuine high-stress days the system catches; precision is the share of its warnings that turn out to be right. The two pull against each other — warn more freely and you catch more, but you are also wrong more often. Where the threshold sits comes down to which error costs more.
Alternatives considered
  • Balance the two and maximise their harmonic mean (F1)
  • Favour precision — warn only when confident, and avoid nagging
  • Favour recall — accept false alarms rather than miss a real one
Trade-off
Favouring recall means more false alarms: a warning that a hard few days may be coming, followed by nothing happening. Repeated often enough that devalues every warning, which is a real cost and not one to wave away.
Decision
Recall, landing at 88%. The two errors are not symmetric: a false alarm costs one unnecessary nudge, a missed one costs a low patch nobody caught. In this domain that asymmetry is large enough to decide the threshold on its own. The wording follows the same logic — the system says something worth keeping an eye on rather than issuing a warning, so a false positive lands gently.

Why self-host inference instead of calling a commercial API?

Situation
Calling a hosted API is by far the easier path: no GPU to run, no model to deploy, and generally better output. But the text being processed is a person's private journal.
Alternatives considered
  • Call a commercial LLM API — fastest to build, least to operate
  • Call a commercial API after de-identifying the text
  • Self-host open-weight models so the data never leaves a controlled environment
Trade-off
Self-hosting is expensive in the ways that matter: a GPU to provide, model serving to build, and outages to own. Open-weight 7B models also generate noticeably weaker Chinese than a large commercial model. De-identification sounds like the middle path, but a journal is identifying by nature — the details of someone's life do not anonymise.
Decision
Self-hosted. Open-weight Qwen2.5-7B-Instruct runs on my own GPU behind FastAPI, reached through a Cloudflare Tunnel. The generation-quality gap against a commercial API is a real price. But 'journal text has never left an environment I control' is a property this product does not get to trade away.
HeartBox system architectureFour layers, top to bottom. The frontend is served from Cloudflare Pages and packaged for Android with Capacitor. It reaches a Django REST Framework backend on Google Cloud Run over REST and WebSocket; that backend also holds Django Channels for realtime messaging and LangChain with ChromaDB for BGE-M3 vector retrieval. The backend calls a self-hosted GPU inference service through a Cloudflare Tunnel — FastAPI running Qwen2.5-7B-Instruct — so journal text never leaves a controlled environment. The backend reads and writes a PostgreSQL database in which journal fields are stored Fernet-encrypted.FrontendCloudflare Pages ・ React ・ Vite ・ TailwindPackaged for Android with CapacitorREST API ・ WebSocketBackend ・ Google Cloud RunDjango REST FrameworkDjango Channels (realtime over WebSocket)LangChain + ChromaDBBGE-M3 embeddings, vector retrievalCloudflare TunnelSelf-hosted GPU inferencellm.heartbox.twFastAPIQwen2.5-7B-InstructOpen-weight ・ handles Traditional ChineseJournal text never leaves this boundaryDatabase ・ PostgreSQLAccounts ・ journal entries ・ remindersJournal fields stored Fernet-encryptedKeys are injected from the environment, held apart from the databaseA database leak alone cannot decrypt journal text
HeartBox system architecture. The dashed box is the self-hosted inference service — the most sensitive text stays inside it.

Why split consent into three separate steps?

Situation
The system collects journal text, the sentiment and stress scores derived from it, and optionally health and sleep data. GDPR defines consent as specific to a purpose (Art. 4(11)), expects distinct processing operations to be separately consentable (Recital 43), and in Art. 7(4) weighs against any service made conditional on consent the contract does not require. Taiwan's Personal Data Protection Act, Art. 7(2), likewise requires consent to out-of-purpose use to be given on its own. A single 'I agree to everything' checkbox does not meet it.
Alternatives considered
  • One consent form covering everything
  • Two stages: terms of service, then data use
  • Three separate stages: data use, AI-training consent, age confirmation
Trade-off
Three steps instead of one click will cost sign-up conversion, and letting people refuse AI-training use means deliberately giving up training data that would otherwise be available. Both losses are real.
Decision
Three separate stages. The load-bearing detail is that AI-training consent is its own checkbox and refusing it leaves every feature available — if refusal degraded the product, the consent would not be freely given in the legal sense. Users aged 13 to 17 go through parental verification: the system emails a confirmation link, and AI features unlock only once it is followed. The conversion cost is what this decision is worth paying.

Why compare users against themselves rather than against a baseline?

Situation
Once the forecast worked, the next question was how to show someone whether they are improving. The obvious answer is to compare against other users — but variation between people is so wide that the comparison carries almost no meaning. A naturally lower baseline mood is not a warning sign.
Alternatives considered
  • Cross-user percentile: 'your score is higher than 70% of users'
  • A fixed absolute threshold: 'below −0.3 counts as bad'
  • A personal baseline: the user's own first week
Trade-off
A personal baseline cannot answer 'how do I compare to other people', and it has nothing to say to a user who just signed up — with too little history the honest output is no output, which is a visible gap in the product experience.
Decision
Personal baseline. Progress is measured against the user's own first week, and an upward arrow always means improvement so nothing has to be memorised. Below fourteen days of history the interface shows how many days remain before the comparison unlocks, rather than computing a number that would not mean anything. This is the decision I am happiest with in the whole project — it admits the data is thin instead of papering over it with something that looks rigorous.

Results

Feature engineering produced 53 features: twelve behavioural and physiological metrics — mean sentiment, mean and peak stress, entry count, sleep duration and quality, deep-sleep proportion, step count, active minutes, HRV, resting heart rate, and bedtime variance — each expanded across 1, 3, 7 and 14-day windows for 48 lag features, plus 5 calendar features such as day of week and journalling streak.

What these numbers show, and what they do not. They are offline cross-validation results: evidence that the model learns something predictive from behavioural history. They are not evidence of clinical validity. There was no controlled trial, and nothing here demonstrates that the warnings improved anyone’s mental health. That distinction matters more than the figures do, and good-looking numbers should not be allowed to blur it.

What I’d do differently

  1. Validate on a temporal split, not random 5-fold. The feature set is dominated by lag variables, so a random split lets information from the same stretch of time land in both training and validation, and the metrics are probably optimistic because of it. Training on earlier periods and predicting later ones is the honest evaluation. This is the project’s biggest methodological weakness and the first thing I would fix.

  2. The −0.4 threshold was never swept. It is a defensible value chosen against real entries, but I never quantified how the false-trigger rate moves at −0.3 or −0.5. That experiment was affordable; I spent the time on features instead.

  3. Retrieval quality itself was never measured. The entire value of the RAG pipeline rests on the retrieved passages being relevant, yet I built no retrieval evaluation set and never measured top-3 hit rate. A model citing a source is not the same as a model citing the right source.

  4. Too many features, diluting the core. A community feed, friends, courses, achievements, breathing exercises — all complete, none of them connected to the two things that actually carry the project. Given the time again I would put it into retrieval evaluation and prospective validation rather than into the hundred-and-third achievement badge.

  5. I underestimated the operational cost of self-hosting. I still think the decision was right, but when the GPU service goes down every AI feature goes with it, and I built no degraded mode. At minimum it should fall back to plain journalling without analysis rather than showing an error.

Stack

LayerTechnology
FrontendReact ・ Vite ・ Tailwind CSS ・ Recharts
BackendDjango REST Framework ・ Django Channels (WebSocket)
AIQwen2.5-7B-Instruct ・ BGE-M3 ・ LangChain ・ ChromaDB
Machine learningscikit-learn (Random Forest)
DatabasePostgreSQL, journal fields Fernet-encrypted (AES-128-CBC + HMAC-SHA256)
MobileCapacitor (Android)
DeploymentCloudflare Pages ・ Cloudflare Tunnel ・ Google Cloud Run

On security and privacy: journal contents are stored under Fernet symmetric encryption with keys injected from the environment and held separately from the database, so a database leak on its own does not recover them. Sign-in supports two-factor authentication and the site is HTTPS throughout. Crisis-keyword detection runs across journalling, AI chat and the anonymous community, and surfaces national helplines immediately when it fires.

Back to projects