The problem

Hong Kong's Leisure and Cultural Services Department publishes sports facility availability through the SmartPLAY system, but finding an available session means navigating the official interface by hand.

The data underneath is also unstable. Sessions open and close constantly as people book, cancel, and release courts, so the useful question is never "what facilities exist" — it is "what is free, where, and when".

Proxying the LCSD API on every request would have been the quickest version of this. It also couples the product directly to a third-party government service: response times become their latency, every visitor generates duplicate upstream load, and an outage on their side becomes an outage on this one.

So the application keeps its own representation instead:

100%Drag to pan

Rendering the diagram…

Users are served from PostgreSQL. The crawler absorbs the unreliability.

What it delivers

An open-source full-stack application for discovering sports facility availability across Hong Kong: scheduled background crawling, availability search filtered by date, district, centre, venue, and facility type, and availability watchers that notify when a session opens.

It also ships webhook delivery in generic, Discord, and Slack formats, English / Traditional Chinese / Simplified Chinese interfaces, anonymous browser sessions instead of accounts, persistent crawl and retry history, and Docker-based deployment with CI/CD.

The visible product is a search interface. The engineering underneath it is the part worth describing.

Architecture

One deployment, clear boundaries: TanStack Start hosting the React frontend and server functions, TanStack server functions bridging into services, services owning Prisma, Prisma owning PostgreSQL. The crawler is a separate workload sharing that database.

100%Drag to pan

Rendering the diagram…

This is deliberately a modular monolith. The crawler, watchers, web UI, and API have different responsibilities, but nothing demonstrated that they needed separate deployments — and splitting them would have bought service discovery, distributed tracing, and inter-service failure modes in exchange.

Designing for an unreliable upstream

The single largest engineering investment in this project was treating the LCSD API as something that will eventually fail, because it eventually does.

Every upstream request sits behind a timeout, a classified retry policy, and a bounded concurrency queue. Timeouts, connection failures, 5xx, and parse failures retry with exponential backoff and jitter; 404 and other 4xx responses are treated as permanent, because retrying a request the upstream has already rejected as invalid only adds load.

Retries cannot run forever, so persistent failures land in a database-backed dead letter queue recording the facility, district, date, error type, HTTP status, attempt count, first and latest failure times, and next retry time. That turns "something failed last night" into a query.

Above both sits a circuit breaker. When LCSD starts failing consistently, continuing to issue requests only wastes application resources and lengthens queues, so the breaker opens, waits out a cooldown, then permits a limited half-open probe before deciding whether to close again.

100%Drag to pan

Rendering the diagram…

Bounded concurrency, not maximum concurrency

The crawler could launch hundreds of requests at once. A bounded p-queue means it processes several in parallel without ever bursting against SmartPLAY.

The external service is a dependency this project does not control, so protecting it is part of protecting itself.

Concurrency and request timing are configurable through environment variables rather than hard-coded, and the crawl schedule runs through node-cron with its own persisted state.

A scheduled run creates individual crawl jobs that are tracked separately, so a multi-day refresh is visible as a set of tracked jobs instead of one opaque process.

Turning an external contract into a domain model

SmartPLAY responses nest periods, districts, venues, facilities, and sessions in a shape that is convenient for SmartPLAY and wrong for everything downstream.

Zod validates the boundary and the crawler normalises it into District, Facility, FacilityType, Session, and CrawlJob rows before storage. After that boundary, the frontend never needs to understand the upstream format — which is what let the UI and the crawler evolve independently.

The database then owns the correctness rules that application code should not have to enforce alone. Sessions carry a composite uniqueness constraint across venue, facility, date, and start time, so repeated crawls update one stable row instead of accumulating duplicates — enforced at the persistence layer, where concurrent writes cannot bypass it.

Indexes follow the actual query shape the booking interface asks: venue + date, date + available, date + facility, and available + verification time.

Availability counts are pre-aggregated into an AvailabilityStats representation across date, district, centre, and facility, so the UI never recomputes counts from every raw session row on a page request.

Availability watchers

A user who cannot find a free session should not have to keep refreshing. A watcher records venue, facility, date, start time, and end time; the crawler refreshes availability; an evaluator compares the observed state and emits a watch hit when it changes.

The notification service delivers that hit as a generic webhook, a Discord payload, or a Slack payload. Notification outcomes are recorded rather than discarded, so a failed delivery is observable instead of silently lost.

Watcher creation is treated as the server-side, abuse-prone operation it is and is protected with Cloudflare Turnstile, alongside security headers and a Content Security Policy configured so Turnstile works without broadly weakening browser security.

Anonymous sessions instead of accounts

Watchers need ownership, but ownership does not require identity. A full registration and password system was not justified by the feature, so the application issues an anonymous browser session that owns a user's watchers and settings, and every ownership check is enforced against that session identifier.

A real authentication system can be added later if a feature genuinely requires persistent user identity. Until then it would only be complexity.

Internationalisation

Hong Kong is not an English-only audience, so localisation is a feature rather than a translation pass at the end.

The application ships English, Traditional Chinese, and Simplified Chinese through i18next and react-i18next, with resources split by namespace, loaded lazily, and converted between Chinese variants where appropriate.

Lifecycle and observability

A crawler produces continuously changing data, so retention is part of the design rather than an afterthought. Cleanup processes remove expired sessions, expired watchers, stale watcher hits, orphaned user settings, dead-letter rows already resolved, and abandoned anonymous browser sessions.

Operating questions are answered from persisted state through Pino structured logging plus crawl job records: did the crawler run, which date failed, which facility fails repeatedly, was the failure retryable, did the webhook delivery succeed. Without those records the same questions require guessing at logs.

Tech stack

  • Frontend: React 19, TanStack Start, TanStack Router, TanStack Query, TanStack Form, Tailwind CSS, Base UI, Lucide
  • Backend: TanStack Start server functions, Node.js, Zod, Pino, node-cron, p-queue
  • Database: PostgreSQL with Prisma ORM
  • Internationalisation: i18next, react-i18next, chinese-conv
  • Security: Cloudflare Turnstile, Content Security Policy, CodeQL
  • Quality: Vitest, Testing Library, Biome, TypeScript
  • Infrastructure: Docker, Docker Compose, GitHub Container Registry, GitHub Actions, Cloudflare Tunnel, Tailscale

Testing and CI/CD

Tests cover crawler operations, scheduler recovery, retry behaviour, session cleanup, utilities, and React components, using a mock crawler repository so the suite never depends on the real SmartPLAY service being up. A government API outage should not be able to turn the build red.

GitHub Actions runs install, Prisma client generation, Biome, TypeScript, and tests before building and publishing an image to GHCR; production deploys only after both quality checks and the image build succeed.

Deployment reaches a private server over Tailscale rather than exposing SSH publicly, pulls the new image, recreates the containers, and finishes with a health check. CodeQL scans pushes to main, pull requests, and a weekly schedule.

What this project demonstrates

The interesting part of SmartPlay HK OSS is not the availability interface — it is the decision to assume the external system will fail, and to answer that with timeout, retry, backoff, jitter, bounded concurrency, circuit breaking, partial-failure recovery, deduplication, observability, and cleanup.

The other lesson was separating the external contract from the internal domain model. SmartPLAY's response format is useful for SmartPLAY and never needed to become the architecture of this application — validating and normalising at the boundary is what let everything downstream work in models that represent the product's own domain.