Files
gacha-event-tracker/README.md
T
Lucas WintherandClaude Opus 5 02863ed008 docs: cover the refresh pipeline, sanitisation and dailies
The status sections claimed the feed was generated from fixtures and the
scheduler unbuilt, which stopped being true. Also documents the
sanitisation stage and the two new key spaces — `dailies:<game>` and
game-day keys — beside the existing warning about event IDs, since they
carry the same "no server-side recovery" property.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
2026-08-15 21:19:08 +02:00

251 lines
12 KiB
Markdown

# Event Clock
Live and upcoming events across your gacha games, sorted by what expires next.
You play three or four gacha games. Each has its own calendar, none of them talk to each other, and
the only question that actually matters — *what runs out first?* — takes four browser tabs to
answer. This does it in one screen.
No account. No login. What you've finished, what you're partway through, how much work you reckon
each event is, and which day of a daily you've ticked off are saved in your browser and never leave
your device.
## Status
Usable, and it now keeps itself up to date. The parsing pipeline and the interface work end to end,
a scheduled job refreshes the sources twice a day, and the whole thing builds, serves, containerises
and deploys. The database and review queue are specified but not built.
| Piece | State |
|---|---|
| Event schema, date parsing, Game8 parser | Built, tested |
| Six game sources | Built, tested |
| Cross-source merge and conflict detection | Built, tested |
| Input sanitization at the ingest boundary | Built, tested |
| Scheduled refresh — robots, snapshots, commit-on-change | Built, tested offline |
| Web interface, daily checklists, offline support | Built |
| Static server, Docker image, GitHub + GitLab CI | Built |
| SQLite, review queue | Specified in `docs/`, not built |
The feed is still a static JSON file rather than a database read: the refresh job commits the raw
pages it fetched, CI rebuilds the feed from them, and a clean checkout with no snapshots falls back
to the checked-in fixtures — so the build stays offline and reproducible either way.
The refresh runner is tested against a fake fetch, never a live wiki. Its first real run against
game8.co and wiki.gg is unproven.
## Try it
```bash
bun install
bun run dev # build, then serve on :3000
```
Then open <http://localhost:3000>.
Or with Docker:
```bash
docker build -t event-clock .
docker run --rm -p 3000:3000 event-clock
```
The image build runs typecheck and tests, and parses only checked-in fixtures — no network, so it is
reproducible and a wiki being down never breaks it.
## Commands
```bash
bun test # full suite, offline, no network
bun run typecheck # tsc --noEmit
bun run build # feed + css + js + html into public/
bun run build:feed # regenerate public/data/events.v1.json from snapshots, else fixtures
# Fetch the sources. Makes real requests, so read "Conduct" first.
bun run refresh --dry-run # plan only: no requests, no writes
bun run refresh --only genshin-game8-events
# Run one source against its fixture and print what it yields
bun run parse genshin-game8-events fixtures/genshin/game8-events-2026-08-14.html
bun run parse nte-game8-events fixtures/nte/game8-events-2026-08-14.html --json
```
## How it works
```
game wikis ─► fetch ─► parse ─► sanitize ─► merge ─► validate ─► gate ─► feed ─► browser
│ │ │ │ │ │
robots, 6h, per-site untrusted per-game hold anything localStorage:
conditional, parser text corroboration uncertain for what you've done,
snapshots bounded and conflicts human review day by day if it
repeats daily
```
**Parsers are deterministic code.** There is no LLM anywhere in the pipeline — no API key, no
inference, no per-run cost. A source that cannot be parsed reliably does not get an adapter, rather
than getting a model that guesses at it.
**Three layers, so sources multiply cheaply.** A *parser* understands one site template (one Game8
parser serves every Game8 page). An *adapter* binds a URL and a game to a parser. *Merge* reconciles
several sources for the same game. Adding a source for a site already covered is one registry entry.
**Nothing is guessed.** Every date function returns null rather than inventing a missing year or
end date. Sources really do publish "July 10, 2026 - Permanent" and "Jul. 24, 2026 - End of 4.6";
those keep their real start and report no end, rendered distinctly from an event ending far away —
never filled in with a plausible-looking date.
That last rule is the whole product. A missing event sends you to a wiki; a confidently wrong end
date makes you miss content. Given the choice, this ships nothing rather than a guess.
## Games
| Game | Source | Events |
|---|---|---|
| Genshin Impact | Game8 | 9 |
| Honkai: Star Rail | Game8 | 6 |
| Wuthering Waves | Game8 | 10 |
| Zenless Zone Zero | Game8 | 12 |
| Neverness to Everness | Game8 | 13 |
| Arknights: Endfield | Game8 + wiki.gg | 6 |
Game8 uses a different page template for almost every game — label/value detail tables, column
tables, rowspan Start/End pairs — and four different date formats between them. One parser handles
all of it; each game costs a registry entry.
Endfield is the first game with two sources. wiki.gg publishes machine-readable ISO timestamps with
one timer per server region, so its events carry exact times and per-region ends — it outranks Game8
and wins when they disagree. Merge caught two real disagreements between them, each 70 hours apart on
the end date; those are flagged rather than averaged.
Its Game8 page yields only two events because most of it genuinely has no dates — every `Duration`
row reads "Permanently Available" and its version grid shows `07/16` with no year. The two dated
events hide in a combined cell (`Period: 08/09/26 - 08/30/26 During the event...`), which is where
the `MM/DD/YY` parser earns its place.
Arknights is defined in the schema and awaiting a source.
## Adding a source
1. Check `robots.txt` and the site's terms. If automated access is forbidden, stop — find another
source.
2. Capture the page once into `fixtures/<game>/<source>-<YYYY-MM-DD>.html`.
3. Reuse an existing parser if the site is already covered; otherwise write one implementing
`SourceParser`.
4. Add an entry to `SOURCES` in `src/ingest/adapters/index.ts`.
5. Write the expected output and a test. Then check a few events against the live page by hand — a
passing test only proves the parser agrees with a file you wrote yourself.
Full walkthrough in `docs/INGESTION.md`, or run the `add-game-source` skill.
## Dailies
Some things are not one job with a deadline. A login campaign is twenty small jobs on twenty
separate deadlines, and a day you miss is gone whatever you do afterwards — which a single "done"
tick cannot express.
So events that repeat get a checklist instead: today's tick, a strip of every day in the run showing
what you got and what you missed, your streak, and how many chances are left. Past days stay
editable, because people tick up later and a checklist you can't correct stops being trusted after
the first mistake.
Alongside them sits **today's dailies** — commissions, sanity, daily training — one tick per game.
No wiki publishes those, so they are a fixed list in the app rather than scraped data, and they are
the only thing on the page that expires tonight rather than next patch.
Both roll over at **04:00 server time** in your region, not midnight, because that is when the games
roll over. Finishing at 02:00 still counts as yesterday.
Repeating events are recognised from what the source actually printed — a login event type, or
wording like "daily", "check-in", "7-day". Nothing is assumed from a game's habits, and an event
whose end was never announced gets a day count rather than a checklist of invented length.
## Sorting
Two orders, and the toggle sits with the list rather than in settings:
- **Ending soonest** — the default, and the reason this app exists.
- **Doing first** — what you're partway through, floated to the top. Ticking a daily counts as
"doing it" without your having to say so twice.
Sorting only ever *groups*. Deadline order survives inside every group, so choosing an order can
never cost you the one thing you came for.
## Offline
The app works with no network. A service worker caches the shell and webfonts, and serves the last
feed it downloaded when the network is gone — countdowns keep running off your device clock. Being
offline is shown in the header and above the footer, because stale data must never look current.
It installs to a home screen as a standalone app.
## Conduct
Sources are community wikis, treated as a guest would: `robots.txt` honoured, a descriptive
`User-Agent` with a contact URL, one request per source per six hours, conditional requests, and raw
snapshots cached so iteration never re-fetches. Every event links back to its source.
`bun run refresh` enforces all of that in code rather than leaving it to good intentions: the
six-hour floor is checked per source, there are no retries (a retry is a second request), and a
`robots.txt` that cannot be read means *do not fetch* rather than *assume yes*. Text scraped from a
page is sanitized at the ingest boundary before it reaches the feed, the browser or your disk.
## Documentation
| Document | Covers |
|---|---|
| `CLAUDE.md` | Working agreements, domain rules, conventions |
| `docs/PRD.md` | What this is, who for, what's out of scope |
| `docs/ARCHITECTURE.md` | Process shape, routes, deployment |
| `docs/DATA-MODEL.md` | Event schema, SQLite tables, client storage |
| `docs/INGESTION.md` | Parser/adapter/merge layers, pipeline stages, review gate |
## CI
Both `.github/workflows/ci.yml` and `.gitlab-ci.yml` run the same gates on every push — typecheck,
tests, and a feed sanity check — then build and publish a container image from the default branch.
GitHub Actions can also deploy to Pages, but that needs two one-time steps it cannot do for itself —
the default `GITHUB_TOKEN` is not allowed to create a Pages site:
1. **Settings → Pages → Source: GitHub Actions.**
2. **Settings → Secrets and variables → Actions → Variables:** add `DEPLOY_PAGES` = `true`.
Until then the `pages` job is skipped and the pipeline stays green. Pages is unavailable for private
repositories on the free plan.
The feed job fails if the event count collapses. A source that quietly stops yielding events is the
failure mode a parser-only pipeline is most prone to, and nothing else would surface it. Tests run
offline against checked-in fixtures, so a red pipeline always means the code changed rather than a
wiki being down.
### Refreshing the data
GitHub Actions only — the GitLab pipeline still runs the gates, but nothing there fetches.
`.github/workflows/refresh.yml` runs `bun run refresh` twice a day (and on demand, with a dry-run
input). It fetches each source at most once per cycle, and **commits only when a page's bytes
actually changed** — a `304`, an identical body, or a fetch that fails to parse all leave the
working tree clean and produce no commit. When something did change it commits the raw snapshots and
dispatches `ci.yml`, which typechecks, tests, rebuilds the feed and deploys through the path that
already existed; none of that logic is duplicated.
A body that yields zero events is rejected and the previous snapshot kept, so a wiki redesign shows
up as a stale timestamp rather than an empty calendar. One source being down is a warning; every
source being down fails the run, so a bad cycle never gets committed.
The schedule is off for forks (it is pinned to this repository) — a fork owner can still dispatch it
by hand and take responsibility for the traffic.
### Hosting under a subpath
Assets resolve against a `<base href>` that the build substitutes, so the app works at a domain root
and under a subpath alike:
```bash
BASE_PATH=/gacha-event-tracker/ bun run build
```
The Pages job sets this automatically. Without it, a subpath deploy 404s on every asset.
## Licence
Not yet chosen. Event data belongs to the sources it came from and is linked back on every event.