The problem
Learning spoken Cantonese is hard for three reasons no existing tool solves well:
- Written resources and spoken language diverge. Most Cantonese material is written Standard Chinese. A learner can read a lyric or a news headline but cannot hear it.
- Romanisation without timing is dead weight. Jyutping tells you what a syllable sounds like, not when it lands in a song. Learners need to see the syllable move with the music.
- Meaning lives in context, not in word lists. Knowing 佢 means "he/she" is useless if you cannot hear it inside real speech.
This platform is built around a single content spine: a large, curated Cantonese lexical corpus — words, characters, pronunciations, meanings, difficulty — plus a lyric corpus that links that vocabulary back to real songs. Everything else is a different way of reading the same spine.
What it does
Area | Capability |
|---|---|
Dictionary | Word and character lookup with Jyutping, Yale and IPA; multiple readings; component characters; synonyms/antonyms; glosses; difficulty and register labels. |
Lyrics | Thousands of songs with searchable lines, artist/lyricist/theme tags, and per-word linking back to dictionary entries. |
Karaoke | Line-by-line and word-by-word singing practice with real timing derived from the audio, so the highlighted word follows the melody. |
Sentences | Themed study sentences with per-word POS, readings, imagery, and translations into four interface languages. |
Talk lessons | Real spoken Cantonese segmented into captioned, translatable, transliterated lines. |
Pronunciation | Tone/initial/final search, an audio reference for every syllable, and speech-assessment exercises. |
Practice loop | Flashcards, echo and speaking challenges, XP/level/streak rewards, and an AI chat partner for translation drills. |
System shape
Three cooperating pieces, deployed together:
- API server — the product's backbone. Serves all learning content, accounts, storage and search from PostgreSQL, and orchestrates the AI and media micro-services.
- Web app — the learner-facing single-process app: a React single-page app plus a thin backend for auth, rewards, analytics and provider proxies. It reads learning content from the API server.
- Micro-services — small, single-purpose services for the language, AI and media work the API server does not want to host.
The web app is the only browser-facing surface. The API server owns all learning content and never exposes a micro-service directly — every request from the browser is answered by the API server, which fans out to the internal services as needed. Internal calls are authenticated, so nothing is callable from outside the deployment.
Core workflows
1. Dictionary lookup
A search resolves to an entry carrying readings, meanings, tags, difficulty and component characters; the entry is then enriched with lyric lines that actually contain the word, and the lexicon page renders from both. A miss returns the standard error envelope rather than an empty page.
That two-part result is the product's central promise working in one place: a learner goes "word → song" instead of stopping at a dictionary entry.
2. Song import to karaoke
The longest pipeline, and the one that produces the karaoke experience.
The app deliberately never invents a boundary. Results carry quality evidence and their provenance, and anything uncertain is surfaced for review instead of silently smoothed over — because an alignment that is plausible but half a second out produces karaoke that looks correct and teaches the wrong thing.
3. Sentence study
A themed sentence is tokenised per word, each word resolved to its part of speech and reading from the lexicon. The sentence is translated, themed and sentiment-labelled, and paired with a scene illustration. The page then shows the sentence's structure, per-word readings, the illustration, the translation and related sentences.
Because annotation is a lookup rather than a hand-written field, adding a sentence does not mean annotating it, and the reading shown in a sentence cannot disagree with the reading shown on the entry — it is the same record.
4. Talk lessons
A recording is segmented on the pauses and phrase structure rather than on fixed intervals, then captioned. An enrichment micro-service adds translation, transliteration, notes and vocabulary; the lesson and its segments are stored; the player offers segmented audio, synchronised captions and per-segment replay.
Transliteration matters here for a specific learner state: someone who cannot read the characters has a translation but no way to connect it to the sound they just heard.
5. Practice and progress
A graded attempt returns a score, feedback and an audio replay. The progress event is recorded idempotently per attempt, so re-attempting does not double-count. XP, level and streak update, and the dashboard reflects the change.
The data pipeline
Content quality is the product, so the data is built deliberately rather than scraped at request time.
Sources are normalised before anything else reads them — public dictionaries disagree about record shape, fields, encoding and character variants, and a source that cannot be normalised is skipped rather than allowed to propagate malformed records. Records are then enriched with romanisation systems beyond the primary one and with every valid reading kept, merged across sources, and classified with difficulty, register and script-variant forms.
The output is a versioned seed snapshot rather than a pile of ad-hoc imports: seeding is repeatable, reruns are idempotent, and a snapshot can be published and consumed by every environment from one place. That is what makes fixing a corpus entry a reviewable pipeline change instead of a manual production database edit.
Architecture
The API server is layered, and the rule is one sentence: domain knows nothing about frameworks, application orchestrates, infrastructure implements the ports, and interface is the only layer that speaks HTTP.
The value of the rule is testability of the part that matters. Business rules are pure TypeScript and framework-free, so they are testable without a database or a request — which is what lets the parts that compute timings and difficulty labels be tested deterministically.
Why each micro-service is separate
Each earns its separation by the same criterion: the work is slow, bursty, or independently failing. Alignment depends on model weights and a Python runtime. Synthesis loads an audio model. Image generation calls an external provider. Enrichment is a batch operation over content. None of these has the same reliability or resource profile as serving a dictionary lookup, and coupling them would mean a model that fails to load also takes down search.
Keeping them small is a second deliberate choice: each is one concern with a contract, so it can be started only when the deployment needs it, scaled independently, and replaced without touching the rest.
Tech stack
Layer | Choice |
|---|---|
API runtime | Deno |
API framework | Hono with schema-driven OpenAPI |
Data access | Drizzle ORM on PostgreSQL, with vector and trigram extensions |
Auth | Better Auth (sessions, OAuth, email) |
Contracts | OpenAPI-first, generated TypeScript clients |
Web app | React with Vite and Tailwind, single-process server |
i18n | Four interface languages |
Storage | S3-compatible object storage |
Observability | OpenTelemetry with structured logs |
Micro-services | Python, one concern each, all behind signed internal APIs |
Deploy | Docker Compose behind a single edge proxy |
Deploy
- The edge proxy is the only public entry point. It routes browser traffic to the web app and API traffic to the API server, and terminates TLS.
- The API server runs with scoped permissions and no general egress — it reaches its database and nothing else.
- Micro-services are internal-only, never published, and each requires a signed request from the API server, so a compromised browser session cannot call a model or download service directly.
- The database sits behind a connection pooler and is not exposed.
- Optional services are started only when the deployment needs them.
Promoting a change is CI-driven: type-check, lint, format, client-drift gates and tests must all pass before an image is built and rolled out.
Quality gates
The gates are organised around places where a silent wrong answer is worse than an error.
- Deterministic API generation — the OpenAPI spec and typed clients are generated, and a drift check fails CI if committed output no longer matches. Committing the generated client is what makes the check meaningful; regenerating during each build would mean nothing ever fails.
- Contract gates — micro-service clients are generated from committed contract snapshots and drift-checked, so a service contract change cannot slip through. A signature verifies the caller, not the shape.
- Schema readiness — startup verifies the schema is present instead of failing later at query time, turning a deployment mismatch into an attributable boot failure.
- Provenance — alignment results carry their quality evidence and the original values, kept auditable rather than repaired into looking correct.
- Tests — co-located unit tests next to the code, plus type, lint and format gates.
License
Private project. Learning content is used for educational purposes; lyrics and recordings remain the property of their rights holders and are used under fair-use for study. Check the repository license before any redistribution.