docs: drop the LLM extraction layer, document multi-source ingestion
Event data is parsed deterministically; there is no model call, API key or per-run cost anywhere in the pipeline. A source that cannot be parsed deterministically gets no adapter rather than an inference fallback. Documents the parser/adapter/merge split, records that Game8 uses three page templates, and adds Neverness to Everness. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 5
parent
493fc9f22a
commit
3ea7286f56
+17
-25
@@ -14,7 +14,7 @@ are no users.
|
|||||||
│ │ │
|
│ │ │
|
||||||
│ ▼ │
|
│ ▼ │
|
||||||
│ ingest pipeline │
|
│ ingest pipeline │
|
||||||
│ fetch → clean → parse|extract → validate │
|
│ fetch → parse → validate │
|
||||||
│ → reconcile → gate → publish │
|
│ → reconcile → gate → publish │
|
||||||
│ │ │ │
|
│ │ │ │
|
||||||
│ │ └──► quarantine │
|
│ │ └──► quarantine │
|
||||||
@@ -35,9 +35,8 @@ are no users.
|
|||||||
filters, region ← never leaves the device
|
filters, region ← never leaves the device
|
||||||
```
|
```
|
||||||
|
|
||||||
Anthropic API calls happen only inside the ingest pipeline. **No request path — not `/api/events`,
|
The pipeline makes no third-party API calls beyond fetching source pages. There is no inference
|
||||||
not a page load — ever calls the model.** If a feature seems to need live inference, it needs a
|
anywhere, at ingest time or in a request path.
|
||||||
precomputed field instead.
|
|
||||||
|
|
||||||
## Layout
|
## Layout
|
||||||
|
|
||||||
@@ -55,24 +54,19 @@ src/
|
|||||||
migrations/ NNN-name.sql, applied in order at boot
|
migrations/ NNN-name.sql, applied in order at boot
|
||||||
queries.ts all SQL lives here — no SQL in route handlers
|
queries.ts all SQL lives here — no SQL in route handlers
|
||||||
ingest/
|
ingest/
|
||||||
scheduler.ts timer + jitter + per-source lock
|
scheduler.ts timer + jitter + per-source lock [not built]
|
||||||
pipeline.ts the 7 stages, orchestration only
|
pipeline.ts the 6 stages, orchestration only [not built]
|
||||||
clean.ts HTML → text reduction before extraction
|
html.ts flat-table HTML reader (no dependency) ✓ built
|
||||||
extract.ts Anthropic client, prompts, batch submission
|
dates.ts deterministic date parsing ✓ built
|
||||||
validate.ts zod parse + calendar sanity rules
|
validate.ts zod parse + calendar sanity rules [not built]
|
||||||
reconcile.ts diff vs published, confidence, conflict detection
|
reconcile.ts diff vs published, confidence, conflicts [not built]
|
||||||
adapters/
|
adapters/
|
||||||
index.ts registry: GameId → Adapter
|
types.ts Adapter interface, ParseContext ✓ built
|
||||||
genshin.ts
|
game8.ts shared Game8 parser (2 table shapes) ✓ built
|
||||||
hsr.ts
|
index.ts registry: adapter id → Adapter ✓ built
|
||||||
zzz.ts
|
|
||||||
wuwa.ts
|
|
||||||
arknights.ts
|
|
||||||
endfield.ts
|
|
||||||
shared/
|
shared/
|
||||||
schema.ts zod schemas — the contract, imported by both sides
|
schema.ts zod schemas — the contract, both sides ✓ built
|
||||||
types.ts z.infer types only
|
time.ts region reset math, duration formatting [not built]
|
||||||
time.ts region reset math, duration formatting
|
|
||||||
client/
|
client/
|
||||||
main.tsx
|
main.tsx
|
||||||
App.tsx
|
App.tsx
|
||||||
@@ -131,21 +125,19 @@ in that area needs an explicit auth story first.
|
|||||||
data — worst case, the game's lane goes stale and gets a warning badge (F7).
|
data — worst case, the game's lane goes stale and gets a warning badge (F7).
|
||||||
- Three consecutive failures for one source raises its `health` to `failing` in `/api/health`. It
|
- Three consecutive failures for one source raises its `health` to `failing` in `/api/health`. It
|
||||||
does not stop the schedule; a wiki being down for a day is normal.
|
does not stop the schedule; a wiki being down for a day is normal.
|
||||||
- Extraction results are written to `extraction_log` with the input hash, so a prompt change can be
|
- Raw snapshots are cached by content hash, so a parser change is always evaluated offline against
|
||||||
evaluated against previously-seen inputs without re-fetching or re-paying.
|
stored pages rather than by re-fetching.
|
||||||
|
|
||||||
## Deployment
|
## Deployment
|
||||||
|
|
||||||
Single process, single SQLite file, no external services beyond the Anthropic API.
|
Single process, single SQLite file, no external services at all.
|
||||||
|
|
||||||
```
|
```
|
||||||
PORT=3000
|
PORT=3000
|
||||||
ADMIN_PORT=3001 # bound to 127.0.0.1
|
ADMIN_PORT=3001 # bound to 127.0.0.1
|
||||||
DATABASE_PATH=./data/events.sqlite
|
DATABASE_PATH=./data/events.sqlite
|
||||||
ANTHROPIC_API_KEY=sk-ant-...
|
|
||||||
INGEST_INTERVAL_MS=21600000
|
INGEST_INTERVAL_MS=21600000
|
||||||
INGEST_ENABLED=true # false for local UI work — never hits the network or the API
|
INGEST_ENABLED=true # false for local UI work — never hits the network or the API
|
||||||
EXTRACTION_MODE=batch # batch | sync
|
|
||||||
CONFIDENCE_THRESHOLD=0.8
|
CONFIDENCE_THRESHOLD=0.8
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
+9
-27
@@ -9,7 +9,7 @@
|
|||||||
import { z } from "zod";
|
import { z } from "zod";
|
||||||
|
|
||||||
export const GameId = z.enum([
|
export const GameId = z.enum([
|
||||||
"genshin", "hsr", "zzz", "wuwa", "arknights", "endfield",
|
"genshin", "hsr", "zzz", "wuwa", "arknights", "endfield", "nte",
|
||||||
]);
|
]);
|
||||||
|
|
||||||
export const EventType = z.enum([
|
export const EventType = z.enum([
|
||||||
@@ -54,7 +54,7 @@ export const GachaEvent = z.object({
|
|||||||
|
|
||||||
status: z.enum(["published", "delisted"]),
|
status: z.enum(["published", "delisted"]),
|
||||||
confidence: z.number().min(0).max(1),
|
confidence: z.number().min(0).max(1),
|
||||||
extractionMethod: z.enum(["parser", "llm", "manual"]),
|
extractionMethod: z.enum(["parser", "manual"]),
|
||||||
|
|
||||||
version: z.number().int().positive(),
|
version: z.number().int().positive(),
|
||||||
firstSeenAt: z.string().datetime(),
|
firstSeenAt: z.string().datetime(),
|
||||||
@@ -78,8 +78,8 @@ true`, with `regionEnds` carrying the three resolved UTC instants. The client pi
|
|||||||
user's stored region (PRD F5). Collapsing these into a single timestamp loses up to 13 hours of
|
user's stored region (PRD F5). Collapsing these into a single timestamp loses up to 13 hours of
|
||||||
accuracy and will make the countdown wrong for two thirds of users.
|
accuracy and will make the countdown wrong for two thirds of users.
|
||||||
|
|
||||||
**`confidence`** is assigned during reconciliation, not by the model's self-report. See
|
**`confidence`** is assigned by the parser and adjusted during merge and reconciliation — see
|
||||||
`docs/LLM-EXTRACTION.md` § Scoring — a model asserting "I am 0.95 confident" is not evidence.
|
`docs/INGESTION.md` § Scoring. It records how firmly the sources pinned the event down.
|
||||||
|
|
||||||
**`status: "delisted"`** means the event stopped appearing at its source. It is never deleted,
|
**`status: "delisted"`** means the event stopped appearing at its source. It is never deleted,
|
||||||
because a source outage would otherwise silently empty the calendar. Delisted events are excluded
|
because a source outage would otherwise silently empty the calendar. Delisted events are excluded
|
||||||
@@ -122,7 +122,7 @@ CREATE TABLE events (
|
|||||||
source_id TEXT NOT NULL,
|
source_id TEXT NOT NULL,
|
||||||
status TEXT NOT NULL DEFAULT 'published',
|
status TEXT NOT NULL DEFAULT 'published',
|
||||||
confidence REAL NOT NULL,
|
confidence REAL NOT NULL,
|
||||||
extraction_method TEXT NOT NULL,
|
extraction_method TEXT NOT NULL, -- 'parser' | 'manual'
|
||||||
version INTEGER NOT NULL DEFAULT 1,
|
version INTEGER NOT NULL DEFAULT 1,
|
||||||
first_seen_at TEXT NOT NULL,
|
first_seen_at TEXT NOT NULL,
|
||||||
updated_at TEXT NOT NULL
|
updated_at TEXT NOT NULL
|
||||||
@@ -147,10 +147,11 @@ CREATE INDEX idx_quarantine_open ON events_quarantine (created_at) WHERE resolve
|
|||||||
|
|
||||||
-- One row per configured source.
|
-- One row per configured source.
|
||||||
CREATE TABLE sources (
|
CREATE TABLE sources (
|
||||||
id TEXT PRIMARY KEY, -- 'genshin-wiki-events'
|
id TEXT PRIMARY KEY, -- '<game>-<site>-<page>', e.g. 'genshin-game8-events'
|
||||||
game TEXT NOT NULL,
|
game TEXT NOT NULL,
|
||||||
url TEXT NOT NULL,
|
url TEXT NOT NULL,
|
||||||
strategy TEXT NOT NULL, -- 'parser' | 'llm' | 'parser_then_llm'
|
parser_id TEXT NOT NULL, -- parser template id, e.g. 'game8'
|
||||||
|
priority INTEGER NOT NULL DEFAULT 0,
|
||||||
min_interval_ms INTEGER NOT NULL DEFAULT 21600000,
|
min_interval_ms INTEGER NOT NULL DEFAULT 21600000,
|
||||||
etag TEXT,
|
etag TEXT,
|
||||||
last_modified TEXT,
|
last_modified TEXT,
|
||||||
@@ -177,26 +178,7 @@ CREATE TABLE ingest_runs (
|
|||||||
events_held INTEGER DEFAULT 0
|
events_held INTEGER DEFAULT 0
|
||||||
);
|
);
|
||||||
|
|
||||||
-- One row per LLM call. Enables replaying prompt changes against past inputs.
|
-- Cached raw snapshots so re-parsing never re-fetches.
|
||||||
CREATE TABLE extraction_log (
|
|
||||||
id TEXT PRIMARY KEY,
|
|
||||||
run_id TEXT NOT NULL REFERENCES ingest_runs(id),
|
|
||||||
source_id TEXT NOT NULL,
|
|
||||||
model TEXT NOT NULL,
|
|
||||||
prompt_version TEXT NOT NULL,
|
|
||||||
input_hash TEXT NOT NULL, -- of the cleaned text sent to the model
|
|
||||||
input_tokens INTEGER,
|
|
||||||
output_tokens INTEGER,
|
|
||||||
cache_read_tokens INTEGER,
|
|
||||||
cache_write_tokens INTEGER,
|
|
||||||
stop_reason TEXT,
|
|
||||||
refusal_category TEXT,
|
|
||||||
duration_ms INTEGER,
|
|
||||||
created_at TEXT NOT NULL
|
|
||||||
);
|
|
||||||
CREATE INDEX idx_extraction_input ON extraction_log (input_hash);
|
|
||||||
|
|
||||||
-- Cached raw + cleaned snapshots so re-extraction never re-fetches.
|
|
||||||
CREATE TABLE snapshots (
|
CREATE TABLE snapshots (
|
||||||
content_hash TEXT PRIMARY KEY,
|
content_hash TEXT PRIMARY KEY,
|
||||||
source_id TEXT NOT NULL,
|
source_id TEXT NOT NULL,
|
||||||
|
|||||||
+180
-138
@@ -1,108 +1,156 @@
|
|||||||
# Ingestion Pipeline
|
# Ingestion Pipeline
|
||||||
|
|
||||||
Seven stages, run per source. Every stage writes its outcome to `ingest_runs` so a failure two days
|
Six stages, run per source. Every stage writes its outcome to `ingest_runs` so a failure two days
|
||||||
ago can be diagnosed without re-running or re-paying.
|
ago can be diagnosed without re-running.
|
||||||
|
|
||||||
```
|
```
|
||||||
fetch → clean → parse|extract → validate → reconcile → gate → publish
|
fetch → parse → merge → validate → reconcile → gate → publish
|
||||||
│
|
│
|
||||||
└──► quarantine
|
└──► quarantine
|
||||||
```
|
```
|
||||||
|
|
||||||
## The adapter contract
|
## No LLM
|
||||||
|
|
||||||
An adapter is the only per-game code. Everything downstream of `parse` is shared.
|
Event data is extracted by deterministic code only. There is no model call anywhere in this
|
||||||
|
pipeline, no API key, and no per-run cost.
|
||||||
|
|
||||||
|
This is a deliberate constraint, not an omission:
|
||||||
|
|
||||||
|
- A source that cannot be parsed deterministically **does not get an adapter.** Report it rather
|
||||||
|
than reaching for inference.
|
||||||
|
- Parser output is reproducible — the same fixture always yields the same events, which is what
|
||||||
|
makes the fixture tests meaningful.
|
||||||
|
- Iterating is free and offline: `bun run parse <adapter-id> <fixture>`.
|
||||||
|
|
||||||
|
If a source's markup is too unstable to parse, the answer is a different source, not a model.
|
||||||
|
|
||||||
|
## Three layers: parsers, adapters, merge
|
||||||
|
|
||||||
|
The layering is what makes a second, third, or tenth source cheap.
|
||||||
|
|
||||||
|
| Layer | Answers | Lives in | Scope |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **Parser** | "How is this *site* laid out?" | `src/ingest/parsers/` | One site template, many games |
|
||||||
|
| **Adapter** | "Which URL, for which game, via which parser?" | `src/ingest/adapters/index.ts` | One page |
|
||||||
|
| **Merge** | "These sources disagree — now what?" | `src/ingest/merge.ts` | One game, many sources |
|
||||||
|
|
||||||
|
Consequences worth internalising:
|
||||||
|
|
||||||
|
- Adding a source for a site already parsed = **one entry in `SOURCES`**. No new parsing code.
|
||||||
|
- Adding a new *site* = one parser module + its `PARSERS` entry, then adapters as above.
|
||||||
|
- A game may have any number of sources. `parseGame(game, documents, now)` runs them all and
|
||||||
|
merges.
|
||||||
|
|
||||||
|
### The parser interface
|
||||||
|
|
||||||
```ts
|
```ts
|
||||||
export interface Adapter {
|
export interface SourceParser {
|
||||||
id: string; // 'genshin-wiki-events'
|
id: string; // "game8"
|
||||||
game: GameId;
|
label: string; // "Game8"
|
||||||
url: string;
|
canParse(html: string): boolean; // structural sanity check
|
||||||
strategy: "parser" | "llm" | "parser_then_llm";
|
parse(html: string, ctx: ParseContext): GachaEvent[];
|
||||||
minIntervalMs?: number; // default 6h
|
|
||||||
|
|
||||||
/** Narrow the cleaned document to just the region containing event data. */
|
|
||||||
select?(cleaned: string): string;
|
|
||||||
|
|
||||||
/**
|
|
||||||
* Deterministic parse. Return null to fall through to LLM extraction
|
|
||||||
* (only meaningful when strategy is 'parser_then_llm').
|
|
||||||
* Pure over its input — no network, no clock, no randomness. This is what
|
|
||||||
* makes fixture tests possible.
|
|
||||||
*/
|
|
||||||
parse?(cleaned: string, ctx: ParseContext): RawEvent[] | null;
|
|
||||||
|
|
||||||
/** Extra instructions appended to the shared extraction prompt. */
|
|
||||||
extractionHints?: string;
|
|
||||||
|
|
||||||
/** Game-specific normalization: reset times, region offsets, patch cadence. */
|
|
||||||
normalize(raw: RawEvent, ctx: ParseContext): GachaEvent;
|
|
||||||
}
|
|
||||||
|
|
||||||
export interface ParseContext {
|
|
||||||
now: string; // injected, never Date.now() — keeps parse pure and testable
|
|
||||||
sourceUrl: string;
|
|
||||||
sourceId: string;
|
|
||||||
game: GameId;
|
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
**`parse` must not read the clock.** It takes `now` from `ctx`. This is what lets a fixture test
|
`canParse` is the redesign tripwire. Without it, a site rewrite makes every selector miss and the
|
||||||
assert exact output for a page captured last March.
|
parser returns zero events — which reads downstream as "this game has no events" rather than as a
|
||||||
|
failure. The adapter throws when `canParse` is false, so the run fails loudly and the previously
|
||||||
|
published events stay put.
|
||||||
|
|
||||||
### Choosing a strategy
|
Keep `canParse` structural, not content-based, and **do not over-fit it**. Game8's own pages differ
|
||||||
|
in attribute quote style (`class="a-table"` on Genshin, `class='a-table'` on NTE), which is exactly
|
||||||
|
the kind of variation a naive check gets wrong. Every regex in `html.ts` is attribute-agnostic for
|
||||||
|
the same reason.
|
||||||
|
|
||||||
| Source shape | Strategy |
|
### The adapter registry
|
||||||
|
|
||||||
|
```ts
|
||||||
|
const SOURCES: SourceSpec[] = [
|
||||||
|
{ id: "genshin-game8-events", game: "genshin",
|
||||||
|
url: "https://game8.co/games/Genshin-Impact/archives/301601", parserId: "game8" },
|
||||||
|
{ id: "nte-game8-events", game: "nte",
|
||||||
|
url: "https://game8.co/games/Neverness-to-Everness/archives/592073", parserId: "game8" },
|
||||||
|
];
|
||||||
|
```
|
||||||
|
|
||||||
|
`priority` (default 0) breaks ties when two sources disagree and neither is clearly better — give
|
||||||
|
official feeds a higher number than community wikis. Adapter ids are `"<game>-<site>-<page>"` and
|
||||||
|
are recorded on every event as `sourceId`, so any row in the feed traces back to the source that
|
||||||
|
produced it.
|
||||||
|
|
||||||
|
### Assessing a new source
|
||||||
|
|
||||||
|
| Source shape | Verdict |
|
||||||
|---|---|
|
|---|---|
|
||||||
| JSON API, or a stable HTML table with consistent headers | `parser` |
|
| JSON API, or an HTML table with consistent headers | Good — write the adapter |
|
||||||
| Free-form patch notes, announcement prose, inconsistent markup | `llm` |
|
| Label/value or column tables with full dates including a year | Good — an existing parser may already handle it |
|
||||||
| Mostly-stable markup that occasionally changes | `parser_then_llm` |
|
| Dates without a year, or no end date at all | **Unsupportable** — yields nothing rather than guessing |
|
||||||
|
| Free-form prose with no table structure | Find a different source |
|
||||||
|
|
||||||
Prefer `parser`. It is free, deterministic, and instantly testable. The LLM exists for sources that
|
Game8 uses at least three page templates and a game's page may use any of them:
|
||||||
genuinely cannot be parsed reliably, not as the default. A source with a clean API that goes through
|
|
||||||
the model is a bug.
|
1. **Label/value detail tables** — `Event Start` / `Event End` rows under a per-event `h3`, full
|
||||||
|
dates with year. *(Genshin Impact)*
|
||||||
|
2. **Column tables** — `Event | Duration | Event Details | Rewards`, one row per event, under a
|
||||||
|
section heading. *(Neverness to Everness)*
|
||||||
|
3. **Image-grid schedules** — a bare `MM/DD`, no year, no end date. **Unsupportable.**
|
||||||
|
*(Arknights: Endfield)*
|
||||||
|
|
||||||
|
Shapes 1 and 2 are handled. Before assuming a new Game8 page will work, dump its heading/table
|
||||||
|
structure and check which shape it uses.
|
||||||
|
|
||||||
## Stage 1 — fetch
|
## Stage 1 — fetch
|
||||||
|
|
||||||
- Send `If-None-Match` / `If-Modified-Since` from `sources.etag` / `last_modified`. A `304` ends
|
- Send `If-None-Match` / `If-Modified-Since` from `sources.etag` / `last_modified`. A `304` ends
|
||||||
the run as `skipped_unchanged` with zero further cost.
|
the run as `skipped_unchanged`.
|
||||||
- `User-Agent: gacha-event-tracker/1.0 (+https://github.com/<owner>/gacha-event-tracker)`.
|
- `User-Agent: gacha-event-tracker/1.0 (+https://github.com/<owner>/gacha-event-tracker)`.
|
||||||
- Honor `robots.txt`. Cache the parsed robots per host for 24h.
|
- Honor `robots.txt`; cache parsed robots per host for 24h.
|
||||||
- 20s timeout; retry twice with exponential backoff on 5xx and network errors; never retry 4xx.
|
- 20s timeout; retry twice with backoff on 5xx and network errors; never retry 4xx.
|
||||||
- Store the raw bytes in `snapshots`.
|
- Store raw bytes in `snapshots`.
|
||||||
|
|
||||||
On failure: increment `consecutive_failures`, leave published events untouched, end the run as
|
On failure: increment `consecutive_failures`, leave published events untouched, end as `failed`. A
|
||||||
`failed`. A source being down never mutates the feed.
|
source being down never mutates the feed.
|
||||||
|
|
||||||
## Stage 2 — clean
|
## Stage 2 — parse
|
||||||
|
|
||||||
Reduce the document before it costs anything. This stage is the second-biggest cost lever after the
|
Hash the raw body (sha256) → `content_hash`. **If it matches `sources.content_hash`, end as
|
||||||
content-hash skip.
|
`skipped_unchanged`** and do no further work.
|
||||||
|
|
||||||
- Drop `<script>`, `<style>`, `<svg>`, `<noscript>`, comments, nav, header, footer, and known
|
Otherwise call `adapter.parse(html, ctx)`, which runs `canParse` and then the parser. Because
|
||||||
wiki chrome (edit links, category boxes, reference lists).
|
parsers are pure, this stage is fully reproducible offline against the stored snapshot:
|
||||||
- Collapse whitespace; convert tables to pipe-delimited text; keep headings as `#` markers so
|
|
||||||
section structure survives.
|
|
||||||
- Apply `adapter.select()` if present to isolate the event region.
|
|
||||||
- Hash the result (sha256) → `content_hash`.
|
|
||||||
|
|
||||||
**If `content_hash` matches `sources.content_hash`, end the run as `skipped_unchanged`.** This is
|
```
|
||||||
the check that keeps a 6-hourly schedule from costing anything on a quiet week — most runs should
|
bun run parse <adapter-id> fixtures/<game>/<source>-<date>.html
|
||||||
end here.
|
```
|
||||||
|
|
||||||
A typical wiki page goes from ~15k tokens raw to ~5k cleaned. Verify with `messages.count_tokens`
|
**Watch the event count.** A source that changes date format or table shape makes events vanish with
|
||||||
when tuning, not by guessing.
|
no error — the parser simply matches nothing. Compare each run's `events_seen` against the previous
|
||||||
|
run and flag a large drop. A source that went from 13 events to 2 has broken, not quieted down.
|
||||||
|
This is the most likely real failure mode of a parser-only pipeline, and nothing else surfaces it.
|
||||||
|
|
||||||
## Stage 3 — parse or extract
|
## Stage 3 — merge
|
||||||
|
|
||||||
Per strategy. `parse` produces `RawEvent[]` directly. `extract` sends the cleaned text to
|
Only meaningful when a game has more than one source; a single-source game passes straight through.
|
||||||
`claude-opus-5` with a structured-output schema — see `docs/LLM-EXTRACTION.md` for the request
|
|
||||||
shape, prompt, and cost rules.
|
|
||||||
|
|
||||||
`parser_then_llm` calls `parse` first and falls through to `extract` only when it returns `null`.
|
`mergeEvents(groups)` compares events across sources:
|
||||||
When that fallthrough happens, log it loudly: it means the source changed shape and the parser needs
|
|
||||||
updating. A `parser_then_llm` source that is silently always falling through is paying LLM prices
|
1. **Same ID** → same event; keep the higher-confidence copy.
|
||||||
for a parser that no longer works.
|
2. **Near match** — same game, title similarity ≥ 0.80, starts within 24h — → same event under
|
||||||
|
different titles; keep the higher-confidence copy.
|
||||||
|
3. **Otherwise** → distinct events; keep both.
|
||||||
|
|
||||||
|
Title similarity alone would merge a rerun with its original, since reruns reuse the name. The
|
||||||
|
start-date proximity check is the actual guard; the title threshold is deliberately loose (0.80) so
|
||||||
|
that "Stygian Onslaught" and "Stygian Onslaught Event" collapse into one row rather than showing
|
||||||
|
the user a duplicate.
|
||||||
|
|
||||||
|
**Agreement raises confidence (+0.10) only across different `sourceId`s.** The same row seen twice
|
||||||
|
in one document is not corroboration.
|
||||||
|
|
||||||
|
**Disagreement is surfaced, never averaged.** Two sources whose `endsAt` differ by more than 24
|
||||||
|
hours produce a `conflicts` entry; the pipeline routes those to quarantine. Splitting the difference
|
||||||
|
between two dates would produce a value neither source asserts — the worst possible answer for a
|
||||||
|
product whose promise is date accuracy.
|
||||||
|
|
||||||
## Stage 4 — validate
|
## Stage 4 — validate
|
||||||
|
|
||||||
@@ -113,109 +161,103 @@ quarantine with `reason: 'sanity_failed'` — never to the feed.
|
|||||||
|
|
||||||
| Rule | Rationale |
|
| Rule | Rationale |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `endsAt > startsAt` when both present | A backwards interval is always a parse error |
|
| `endsAt` after `startsAt` when both present | A backwards interval is always a parse error |
|
||||||
| Duration ≤ 180 days | Patch cycles are ~6 weeks; 180d means a year was misread as a range |
|
| Duration under 180 days | Patch cycles are ~6 weeks; longer means a misread year |
|
||||||
| `startsAt` within [now − 2y, now + 1y] | Catches century typos and relative-date misreads |
|
| `startsAt` within [now − 2y, now + 1y] | Catches century typos and relative-date misreads |
|
||||||
| `endsAt` null ⟺ `endPrecision === "unknown"` | The two fields must agree |
|
| `endsAt` null exactly when `endPrecision` is `"unknown"` | The two fields must agree |
|
||||||
| `regionEnds` non-null ⟺ `regionScoped` | Same |
|
| `regionEnds` non-null exactly when `regionScoped` | Same |
|
||||||
| All `regionEnds` values within 24h of each other | Region resets differ by hours, not days |
|
| All `regionEnds` values within 24h of each other | Region resets differ by hours, not days |
|
||||||
| `title` non-empty, ≤ 200 chars, not a placeholder ("TBD", "Event", "Unknown") | Catches header rows scraped as events |
|
| `title` non-empty, ≤ 200 chars, not a placeholder | Catches header rows scraped as events |
|
||||||
|
|
||||||
|
Rules 1, 4, and 5 are enforced by `GachaEvent` itself in `src/shared/schema.ts`, so they cannot be
|
||||||
|
bypassed by constructing an event object directly.
|
||||||
|
|
||||||
**Soft rules (reduce confidence, do not reject):**
|
**Soft rules (reduce confidence, do not reject):**
|
||||||
|
|
||||||
- Duration under 1 hour or over 60 days → −0.2
|
- Duration under 1 hour or over 60 days → −0.2
|
||||||
- `startPrecision` or `endPrecision` is `"day"` → −0.1
|
|
||||||
- Title very similar to another event in the same batch → −0.15 (likely a duplicate row)
|
- Title very similar to another event in the same batch → −0.15 (likely a duplicate row)
|
||||||
|
|
||||||
## Stage 5 — reconcile
|
## Stage 5 — reconcile
|
||||||
|
|
||||||
Diff the validated candidates against currently published events for this source.
|
Diff validated candidates against currently published events.
|
||||||
|
|
||||||
1. **Exact ID match** → compare fields. Unchanged: no-op. Changed: candidate update.
|
1. **Exact ID match** → compare fields. Unchanged: no-op. Changed: update.
|
||||||
2. **Near match** — same game, date windows overlap, title similarity ≥ 0.85 — → treat as an update
|
2. **Near match** → update the existing event, **keeping the existing ID**. This is what survives a
|
||||||
to the existing event, **keeping the existing ID**. This is what survives a wiki renaming an
|
wiki renaming an event without orphaning every user's completion mark.
|
||||||
event without orphaning every user's completion mark.
|
|
||||||
3. **No match** → new event.
|
3. **No match** → new event.
|
||||||
4. **Published event absent from this run's candidates** → mark `status = 'delisted'`. Do not
|
4. **Published event absent from this run** → mark `status = 'delisted'`. Never delete.
|
||||||
delete.
|
|
||||||
|
|
||||||
**Conflict detection.** A candidate that changes an already-published `endsAt` by more than 24 hours
|
**Conflict detection.** A candidate moving an already-published `endsAt` by more than 24 hours is a
|
||||||
is a `date_conflict`. This is the case worth being paranoid about: the user may have already planned
|
`date_conflict`. The user may have planned around the old date, so route it to quarantine regardless
|
||||||
around the old date, and a silent shift is exactly the failure the product exists to prevent. Route
|
of confidence.
|
||||||
it to quarantine regardless of confidence.
|
|
||||||
|
|
||||||
### Scoring
|
### Scoring
|
||||||
|
|
||||||
Confidence is computed here, from evidence — **not** taken from the model's self-report. A model
|
Confidence records how firmly the sources pinned an event down, so the gate can hold back weak
|
||||||
saying "confidence: 0.95" is a token prediction, not a measurement.
|
cases. The parser assigns a base score; merge and reconcile adjust it.
|
||||||
|
|
||||||
```
|
```
|
||||||
base parser → 0.95 llm → 0.70
|
base 0.95
|
||||||
+0.15 the same event was extracted identically from a previous run
|
−0.05 a boundary is day-precision rather than exact
|
||||||
+0.10 both timestamps have precision "exact"
|
−0.15 the end date is unknown (endsAt null)
|
||||||
+0.10 a second source for the same game corroborates within 1 hour
|
+0.10 an independent source corroborates
|
||||||
|
+0.15 identical event parsed in a previous run
|
||||||
−0.20 any soft rule fired
|
−0.20 any soft rule fired
|
||||||
−0.30 this is a date_conflict against a published event
|
−0.30 a date_conflict against a published event
|
||||||
```
|
```
|
||||||
|
|
||||||
Clamp to [0, 1]. `CONFIDENCE_THRESHOLD` (default 0.8) is the gate.
|
Clamp to [0, 1]. `CONFIDENCE_THRESHOLD` (default 0.8) is the gate. Under the current parser a
|
||||||
|
day-precision event with a known end scores 0.85 and publishes, while one with an unknown end
|
||||||
|
scores 0.75 and is held — the intended bias.
|
||||||
|
|
||||||
## Stage 6 — gate
|
## Stage 6 — gate and publish
|
||||||
|
|
||||||
| Condition | Destination |
|
| Condition | Destination |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `confidence >= CONFIDENCE_THRESHOLD` and no conflict | publish |
|
| Confidence at or above threshold, no conflict | publish |
|
||||||
| `confidence < CONFIDENCE_THRESHOLD` | quarantine — `low_confidence` |
|
| Confidence below threshold | quarantine, `low_confidence` |
|
||||||
| `date_conflict` | quarantine — `date_conflict`, with `conflicts_with` set |
|
| Cross-source or cross-run date disagreement | quarantine, `date_conflict` |
|
||||||
| failed a hard rule | quarantine — `sanity_failed` |
|
| Failed a hard rule | quarantine, `sanity_failed` |
|
||||||
| new event type or field the schema does not recognize | quarantine — `novel_shape` |
|
| Shape the schema does not recognise | quarantine, `novel_shape` |
|
||||||
|
|
||||||
A quarantined event does not block its siblings. If eight events in a run pass and two are held, the
|
A quarantined event does not block its siblings — if eight pass and two are held, the eight publish.
|
||||||
eight publish.
|
|
||||||
|
|
||||||
## Stage 7 — publish
|
Publish upserts by ID in a transaction. Bump `version` and `updatedAt` only when a field actually
|
||||||
|
changed, or the freshness badge (PRD F7) becomes meaningless. Update `sources.content_hash`,
|
||||||
Upsert by ID inside a transaction. Bump `version` and `updatedAt` only when a field actually
|
`etag`, `last_success_at`, and reset `consecutive_failures`.
|
||||||
changed — an unchanged run must not churn `updatedAt`, or the freshness badge (PRD F7) becomes
|
|
||||||
meaningless. Update `sources.content_hash`, `etag`, `last_success_at`, and reset
|
|
||||||
`consecutive_failures`.
|
|
||||||
|
|
||||||
## The review gate
|
## The review gate
|
||||||
|
|
||||||
Quarantined events surface at `GET /review` on the admin listener (`127.0.0.1:ADMIN_PORT`). See
|
Quarantined events surface at `GET /review` on the admin listener (`127.0.0.1:ADMIN_PORT`). See
|
||||||
`docs/ARCHITECTURE.md` § Why `/review` needs no auth — the routes are simply not registered on the
|
`docs/ARCHITECTURE.md` § Why `/review` needs no auth.
|
||||||
public listener.
|
|
||||||
|
|
||||||
The review UI shows, per held event: the parsed fields, the reason and detail, the conflicting
|
- `POST /api/review/:id/approve` — writes to `events` with `extraction_method: 'manual'`,
|
||||||
published event side-by-side when applicable, a link to the source, and the exact cleaned text
|
`confidence: 1.0`. Approving with edits is supported; the corrected value publishes.
|
||||||
excerpt the extraction came from. A reviewer needs to answer "is this date right?" without leaving
|
- `POST /api/review/:id/reject` — stamps resolution only. The candidate is held again next run if
|
||||||
the page.
|
the source has not changed, which is intended.
|
||||||
|
|
||||||
- `POST /api/review/:id/approve` — writes to `events` with `extraction_method: 'manual'` and
|
**Quarantine depth is the pipeline's health signal.** A growing queue means a source changed shape.
|
||||||
`confidence: 1.0`, and stamps `resolved_at` / `resolution`.
|
`/api/health` exposes the count.
|
||||||
- `POST /api/review/:id/reject` — stamps resolution only. The event is not published, and the same
|
|
||||||
candidate will be re-held on the next run if the source has not changed.
|
|
||||||
|
|
||||||
Approving with edits is supported: the reviewer can correct a date before approving. That corrected
|
|
||||||
value is the one that publishes.
|
|
||||||
|
|
||||||
**Quarantine depth is the health signal for the whole pipeline.** A growing queue means a source
|
|
||||||
changed shape or the prompt regressed. `/api/health` exposes the count; watch it.
|
|
||||||
|
|
||||||
## Testing
|
## Testing
|
||||||
|
|
||||||
Every adapter ships:
|
Every adapter ships:
|
||||||
|
|
||||||
1. `fixtures/<game>/<source>-<YYYY-MM-DD>.html` — a real captured page.
|
1. `fixtures/<game>/<source>-<YYYY-MM-DD>.html` — a real captured page.
|
||||||
2. `fixtures/<game>/<source>-<YYYY-MM-DD>.expected.json` — the exact `GachaEvent[]` it should
|
2. `fixtures/<game>/<source>-<YYYY-MM-DD>.expected.json` — the exact `GachaEvent[]` it produces.
|
||||||
produce.
|
3. A test running `parse` against the fixture with a pinned `ctx.now`, asserting deep equality.
|
||||||
3. A test running `parse` + `normalize` against the fixture with a pinned `ctx.now`, asserting
|
|
||||||
deep equality.
|
|
||||||
|
|
||||||
`bun test` must pass with no network access. When a source changes shape, capture a new fixture
|
`bun test` must pass with no network access.
|
||||||
alongside the old one and keep both — the old fixture is the regression test proving the parser
|
|
||||||
still handles the previous format.
|
|
||||||
|
|
||||||
Validator rules get their own unit tests with deliberately broken inputs: backwards intervals,
|
**Regenerating `.expected.json` from the parser makes the test self-consistent, not correct.** After
|
||||||
1000-year durations, `endsAt` set with `endPrecision: "unknown"`. These rules are the last line of
|
an intentional change, re-verify a sample against the live page — and ideally extract the same data
|
||||||
defense before a wrong date reaches a user; test them like it.
|
a second way (a throwaway script over the fixture) to confirm counts and dates independently. That
|
||||||
|
independent check is what caught the exact event counts for both current adapters.
|
||||||
|
|
||||||
|
When a source changes shape, capture a new fixture **alongside** the old one and keep both — the old
|
||||||
|
fixture is the regression test proving the parser still handles the previous format.
|
||||||
|
|
||||||
|
`test/dates.test.ts` covers the cases that matter most: a missing year returns null rather than
|
||||||
|
guessing, impossible calendar dates are rejected, ranges crossing New Year roll the start year back,
|
||||||
|
and abbreviated months parse. `test/merge.test.ts` covers cross-source agreement, disagreement, and
|
||||||
|
rerun disambiguation. These are the last line of defense before a wrong date reaches a user.
|
||||||
|
|||||||
@@ -1,291 +0,0 @@
|
|||||||
# LLM Extraction
|
|
||||||
|
|
||||||
Stage 3 of the pipeline, for sources whose markup is too unstable to parse deterministically. Read
|
|
||||||
`docs/INGESTION.md` first for where this sits.
|
|
||||||
|
|
||||||
## Rules
|
|
||||||
|
|
||||||
1. **The model runs at ingestion time only.** No request path calls Anthropic. A page load must
|
|
||||||
never trigger inference.
|
|
||||||
2. **Deterministic parsers win.** If a source has a JSON API or a stable table, it gets a parser,
|
|
||||||
not a prompt.
|
|
||||||
3. **The model's job is transcription, not judgment.** It converts prose and tables into structured
|
|
||||||
dates. It does not decide what is important, does not infer missing dates, and does not resolve
|
|
||||||
contradictions — it reports them.
|
|
||||||
4. **Confidence is computed from evidence in `reconcile`, not asserted by the model.** The output
|
|
||||||
schema has no confidence field. A model claiming 0.95 confidence has predicted a token, not
|
|
||||||
measured anything.
|
|
||||||
|
|
||||||
## Model and parameters
|
|
||||||
|
|
||||||
| Setting | Value | Why |
|
|
||||||
|---|---|---|
|
|
||||||
| `model` | `claude-opus-5` | Exact, complete ID — never append a date suffix |
|
|
||||||
| `max_tokens` | `16000` | Non-streaming; keeps the request under SDK HTTP timeouts |
|
|
||||||
| `output_config.effort` | `"medium"` | Transcription, not reasoning. Sweep low/medium/high against fixtures before settling |
|
|
||||||
| `output_config.format` | `zodOutputFormat(ExtractionResult)` | Schema-conformant output, validated by the SDK |
|
|
||||||
| `thinking` | *omit* | On by default on `claude-opus-5`; the default is correct here |
|
|
||||||
| `temperature` / `top_p` / `top_k` | **never set** | Removed on `claude-opus-5` — sending any of them returns 400 |
|
|
||||||
|
|
||||||
`thinking: {type: "enabled", budget_tokens: N}` is also removed and returns 400. If you want less
|
|
||||||
thinking, lower `effort`.
|
|
||||||
|
|
||||||
Note that `max_tokens` caps thinking *and* output together. If extraction on a large page returns
|
|
||||||
`stop_reason: "max_tokens"`, raise it rather than trimming the schema.
|
|
||||||
|
|
||||||
## The output schema
|
|
||||||
|
|
||||||
The model returns a list of candidate events plus explicit uncertainty. Note what is **absent**:
|
|
||||||
no confidence score, no "importance" ranking, no summary of the page.
|
|
||||||
|
|
||||||
```ts
|
|
||||||
// src/ingest/extract.ts
|
|
||||||
import { z } from "zod";
|
|
||||||
|
|
||||||
const ExtractedEvent = z.object({
|
|
||||||
title: z.string().describe("The event name exactly as written in the source, not paraphrased."),
|
|
||||||
type: z.enum(["banner","story","rerun","challenge","login","shop","maintenance","other"]),
|
|
||||||
summary: z.string().nullable()
|
|
||||||
.describe("One sentence from the source describing the event. Null if the source gives none."),
|
|
||||||
|
|
||||||
startsAt: z.string().nullable()
|
|
||||||
.describe("ISO 8601 UTC. Null only if the source truly does not state a start."),
|
|
||||||
startPrecision: z.enum(["exact","day","unknown"])
|
|
||||||
.describe("'exact' if a time of day is stated; 'day' if only a date; 'unknown' if neither."),
|
|
||||||
|
|
||||||
endsAt: z.string().nullable()
|
|
||||||
.describe("ISO 8601 UTC. Null when the source says TBD, 'until further notice', or gives no end."),
|
|
||||||
endPrecision: z.enum(["exact","day","unknown"]),
|
|
||||||
|
|
||||||
regionScoped: z.boolean()
|
|
||||||
.describe("True if the end follows each server region's daily reset rather than one global instant."),
|
|
||||||
|
|
||||||
sourceTimezone: z.string().nullable()
|
|
||||||
.describe("The timezone the source stated, e.g. 'UTC+8', 'server local'. Null if unstated."),
|
|
||||||
evidence: z.string()
|
|
||||||
.describe("The verbatim span from the input that gave you the dates. Must appear in the input."),
|
|
||||||
});
|
|
||||||
|
|
||||||
const ExtractionResult = z.object({
|
|
||||||
events: z.array(ExtractedEvent),
|
|
||||||
ambiguities: z.array(z.object({
|
|
||||||
title: z.string(),
|
|
||||||
issue: z.string().describe("What is unclear or contradictory in the source."),
|
|
||||||
})).describe("Events you could not confidently transcribe. These are held for human review."),
|
|
||||||
});
|
|
||||||
```
|
|
||||||
|
|
||||||
`evidence` is the load-bearing field. It is checked in `validate`: if the quoted span does not
|
|
||||||
appear in the input text, the event is quarantined as `sanity_failed`. That check is what turns a
|
|
||||||
fabricated date into a caught error instead of a shipped one.
|
|
||||||
|
|
||||||
## The system prompt
|
|
||||||
|
|
||||||
Kept in `src/ingest/prompts/extract-events.v1.md`, versioned in the filename, and logged as
|
|
||||||
`prompt_version` in `extraction_log` so a regression can be traced to a specific revision.
|
|
||||||
|
|
||||||
It must stay **above 512 tokens** — that is the minimum cacheable prefix on `claude-opus-5`. Below
|
|
||||||
it, prompt caching silently stops working with no error. After any prompt edit, check
|
|
||||||
`usage.cache_read_input_tokens` is non-zero on the second call.
|
|
||||||
|
|
||||||
```markdown
|
|
||||||
You extract scheduled in-game events from gacha game source pages into structured data.
|
|
||||||
|
|
||||||
Your output is consumed by a calendar that players rely on to avoid missing limited-time content.
|
|
||||||
A wrong end date is worse than a missing event: a missing event sends someone to a wiki, a wrong
|
|
||||||
one makes them miss content permanently. Transcribe what the source says; never supply what it
|
|
||||||
omits.
|
|
||||||
|
|
||||||
## What counts as an event
|
|
||||||
|
|
||||||
Anything with a start and a bounded or open-ended run: character and weapon banners, story
|
|
||||||
chapters, side events, login campaigns, limited shops, combat cycles, announced maintenance.
|
|
||||||
|
|
||||||
Not events: permanent features, general game descriptions, patch version numbers on their own,
|
|
||||||
speculation or leaks, community posts, and anything phrased as expected, rumored, or datamined.
|
|
||||||
|
|
||||||
## Dates
|
|
||||||
|
|
||||||
- Emit UTC ISO 8601 with an explicit `Z`.
|
|
||||||
- When the source states a timezone (commonly UTC+8 for Chinese-developed titles), convert to UTC
|
|
||||||
and record what it stated in `sourceTimezone`.
|
|
||||||
- When only a date is given, set the timestamp to 00:00:00Z and `precision: "day"`. Do not guess a
|
|
||||||
time of day.
|
|
||||||
- When the source says the end is TBD, "until further notice", "with the next version update", or
|
|
||||||
gives no end at all: `endsAt: null` and `endPrecision: "unknown"`. This is a correct, expected
|
|
||||||
answer. Do not compute a plausible date from a typical patch length.
|
|
||||||
- `regionScoped` is true when the end is tied to each server's daily reset, false when the source
|
|
||||||
gives one simultaneous global instant. Character banners are usually global; story and login
|
|
||||||
events are usually region-scoped. Use what the source says over this heuristic when it says
|
|
||||||
anything.
|
|
||||||
|
|
||||||
## Evidence
|
|
||||||
|
|
||||||
For every event, `evidence` must be a verbatim span copied from the input that contains the dates
|
|
||||||
you reported. It is checked against the input automatically. If you cannot quote a span, the event
|
|
||||||
belongs in `ambiguities` instead.
|
|
||||||
|
|
||||||
## Ambiguities
|
|
||||||
|
|
||||||
Put an entry in `ambiguities` — not in `events` — when the source contradicts itself, gives dates
|
|
||||||
you cannot reconcile, or describes something that may not be a scheduled event. A human reviews
|
|
||||||
these. Reporting uncertainty is a successful outcome, not a failure; guessing to avoid it is the
|
|
||||||
one thing that breaks this system.
|
|
||||||
|
|
||||||
## Scope
|
|
||||||
|
|
||||||
Report every qualifying event on the page and nothing else. Do not rank them, do not summarize the
|
|
||||||
page, do not comment on your process, and do not add fields the schema does not have. If the page
|
|
||||||
contains no events, return empty arrays.
|
|
||||||
```
|
|
||||||
|
|
||||||
### Why the prompt reads the way it does
|
|
||||||
|
|
||||||
`claude-opus-5` follows instructions literally and verifies its own work without being told, so the
|
|
||||||
prompt states scope and boundaries plainly instead of adding emphasis or self-check scaffolding.
|
|
||||||
Specifically: **do not add "double-check your answer" or "verify before responding"** here. On this
|
|
||||||
model that produces over-verification with no accuracy gain. If extraction quality drops, change the
|
|
||||||
schema descriptions or `effort` — not the volume of the prompt.
|
|
||||||
|
|
||||||
## Request shape
|
|
||||||
|
|
||||||
```ts
|
|
||||||
import Anthropic from "@anthropic-ai/sdk";
|
|
||||||
import { zodOutputFormat } from "@anthropic-ai/sdk/helpers/zod";
|
|
||||||
|
|
||||||
const client = new Anthropic(); // reads ANTHROPIC_API_KEY from env
|
|
||||||
|
|
||||||
const response = await client.messages.parse({
|
|
||||||
model: "claude-opus-5",
|
|
||||||
max_tokens: 16000,
|
|
||||||
system: [
|
|
||||||
{
|
|
||||||
type: "text",
|
|
||||||
text: EXTRACTION_SYSTEM_PROMPT, // stable across every source — cached
|
|
||||||
cache_control: { type: "ephemeral", ttl: "1h" },
|
|
||||||
},
|
|
||||||
],
|
|
||||||
output_config: {
|
|
||||||
effort: "medium",
|
|
||||||
format: zodOutputFormat(ExtractionResult),
|
|
||||||
},
|
|
||||||
messages: [
|
|
||||||
{
|
|
||||||
role: "user",
|
|
||||||
content: [
|
|
||||||
`Game: ${adapter.game}`,
|
|
||||||
`Source: ${adapter.url}`,
|
|
||||||
`Today (UTC): ${ctx.now}`,
|
|
||||||
adapter.extractionHints ?? "",
|
|
||||||
"",
|
|
||||||
"--- SOURCE TEXT ---",
|
|
||||||
cleanedText,
|
|
||||||
].join("\n"),
|
|
||||||
},
|
|
||||||
],
|
|
||||||
});
|
|
||||||
|
|
||||||
// parsed_output is null if the model refused or hit max_tokens — check before use.
|
|
||||||
const result = response.parsed_output;
|
|
||||||
```
|
|
||||||
|
|
||||||
**Ordering matters for caching.** The system prompt is byte-identical across every source, so it
|
|
||||||
sits first and stays cached. Everything volatile — game, URL, `now`, the page text — goes in the
|
|
||||||
user turn, after the cache breakpoint. Interpolating `ctx.now` into the system prompt would
|
|
||||||
invalidate the cache on every single call; it is in the user turn for exactly that reason.
|
|
||||||
|
|
||||||
## Batch mode
|
|
||||||
|
|
||||||
`EXTRACTION_MODE=batch` is the default for scheduled runs: 50% cheaper, results typically within an
|
|
||||||
hour, which is irrelevant against a 6-hour cadence. The synchronous path above is for
|
|
||||||
`refresh --now` and for local development.
|
|
||||||
|
|
||||||
```ts
|
|
||||||
const batch = await client.messages.batches.create({
|
|
||||||
requests: sourcesToExtract.map((s) => ({
|
|
||||||
custom_id: s.runId, // key results by this — order is not guaranteed
|
|
||||||
params: { /* same body as above */ },
|
|
||||||
})),
|
|
||||||
});
|
|
||||||
|
|
||||||
// Poll batches.retrieve(batch.id) until processing_status === "ended",
|
|
||||||
// then stream batches.results(batch.id) and key each result by custom_id.
|
|
||||||
```
|
|
||||||
|
|
||||||
Results arrive in **any order**. Key by `custom_id`, never by array position. The batch job's
|
|
||||||
poll loop lives in `scheduler.ts` and persists `batch.id` so a process restart resumes rather than
|
|
||||||
resubmitting.
|
|
||||||
|
|
||||||
## Handling non-success responses
|
|
||||||
|
|
||||||
Check `stop_reason` before touching `parsed_output`:
|
|
||||||
|
|
||||||
| `stop_reason` | Meaning | Action |
|
|
||||||
|---|---|---|
|
|
||||||
| `end_turn` | Normal | Proceed |
|
|
||||||
| `max_tokens` | Output truncated; `parsed_output` unusable | Raise `max_tokens`, or split the page by section. Fail the run — never publish a partial list |
|
|
||||||
| `refusal` | Safety classifier declined | Log `stop_details.category` to `extraction_log.refusal_category`, fail the run, alert. Vanishingly unlikely for game wiki content — if it fires, the input is probably not what you think it is |
|
|
||||||
|
|
||||||
`stop_details` can be `null` even on a refusal, so branch on `stop_reason` and treat `stop_details`
|
|
||||||
as informational.
|
|
||||||
|
|
||||||
**Do not write a JSON-repair or regex-extraction fallback.** Structured outputs guarantee schema
|
|
||||||
conformance; if parsing fails, the response was truncated or refused, and both cases are handled
|
|
||||||
above. A repair path would silently paper over truncation and publish half a calendar.
|
|
||||||
|
|
||||||
## Cost
|
|
||||||
|
|
||||||
At `claude-opus-5` rates — $5/MTok input, $25/MTok output.
|
|
||||||
|
|
||||||
One extraction of a cleaned wiki page:
|
|
||||||
|
|
||||||
```
|
|
||||||
input ~5,000 tok × $5/MTok = $0.025
|
|
||||||
output ~1,500 tok × $25/MTok = $0.0375
|
|
||||||
───────
|
|
||||||
sync $0.063
|
|
||||||
batch $0.031 (50% off)
|
|
||||||
```
|
|
||||||
|
|
||||||
Six sources, four runs a day:
|
|
||||||
|
|
||||||
| Scenario | Extractions/day | Cost/day | Cost/month |
|
|
||||||
|---|---|---|---|
|
|
||||||
| Realistic — ~20% of runs see changed content | ~5 | $0.16 | **~$4.70** |
|
|
||||||
| Worst case — every source changes every run | 24 | $0.74 | ~$22 |
|
|
||||||
| No content-hash skip, no batch | 24 | $1.51 | ~$45 |
|
|
||||||
|
|
||||||
The gap between rows one and three is the whole argument for the skip check and batch mode. Track
|
|
||||||
actuals by summing `extraction_log` token columns — do not rely on these estimates once there is
|
|
||||||
real data.
|
|
||||||
|
|
||||||
Prompt caching contributes modestly (the ~1k-token system prompt at 0.1× on reads) but is free to
|
|
||||||
keep. Its real value is that it makes the prompt cheap to grow if extraction quality needs more
|
|
||||||
instruction.
|
|
||||||
|
|
||||||
## Evaluating a prompt change
|
|
||||||
|
|
||||||
`extraction_log` stores `input_hash` for every call, and `snapshots` stores the cleaned text keyed
|
|
||||||
by the same hash. So a prompt revision is evaluated by replaying past inputs — no re-fetching, no
|
|
||||||
new scraping load:
|
|
||||||
|
|
||||||
1. Pull the last N distinct `input_hash` values with known-correct expected output.
|
|
||||||
2. Run the new prompt against each cleaned snapshot.
|
|
||||||
3. Diff against `fixtures/*/*.expected.json`.
|
|
||||||
4. Compare on three axes: dates correct, events missed, events hallucinated. **Hallucinated events
|
|
||||||
and wrong dates are disqualifying; a missed event is a regression to weigh.** That asymmetry is
|
|
||||||
the product rule from `docs/PRD.md` restated as an eval criterion.
|
|
||||||
|
|
||||||
Bump the prompt filename version and record it as `prompt_version` so the log distinguishes
|
|
||||||
"extraction got worse" from "the source changed".
|
|
||||||
|
|
||||||
## Things not to do
|
|
||||||
|
|
||||||
- Do not ask the model to output a confidence score. Confidence is computed in `reconcile`.
|
|
||||||
- Do not ask the model to resolve a contradiction between two sources. Route it to quarantine.
|
|
||||||
- Do not send raw HTML. Always clean first — it is a 3× cost difference and it improves accuracy.
|
|
||||||
- Do not add a second model call to "verify" the first. That is a scaffolding pattern this model
|
|
||||||
does not need, and it doubles cost for no measured gain. The evidence-span check and the
|
|
||||||
validator rules are the verification layer.
|
|
||||||
- Do not lower `max_tokens` to save money. Output tokens scale with the number of events found;
|
|
||||||
truncation costs a whole run.
|
|
||||||
+7
-5
@@ -41,9 +41,11 @@ event, because a missing event sends them to a wiki while a wrong one makes them
|
|||||||
| Wuthering Waves | `wuwa` |
|
| Wuthering Waves | `wuwa` |
|
||||||
| Arknights | `arknights` |
|
| Arknights | `arknights` |
|
||||||
| Arknights: Endfield | `endfield` |
|
| Arknights: Endfield | `endfield` |
|
||||||
|
| Neverness to Everness | `nte` |
|
||||||
|
|
||||||
Adding a seventh game must require no schema change — only a new adapter. That is the test of
|
Adding a game must require no schema change — only a `GameId` entry and a source registration.
|
||||||
whether the data model is right.
|
That is the test of whether the data model is right. A game may have several sources; see
|
||||||
|
`docs/INGESTION.md` § Three layers.
|
||||||
|
|
||||||
### Features
|
### Features
|
||||||
|
|
||||||
@@ -106,8 +108,8 @@ The app's entire value is that the dates are right. Therefore:
|
|||||||
|
|
||||||
- An event with an uncertain end date is published with `endsAt: null`, **not** with a plausible
|
- An event with an uncertain end date is published with `endsAt: null`, **not** with a plausible
|
||||||
guess.
|
guess.
|
||||||
- An event whose extraction confidence is below threshold is not published at all until a human
|
- An event whose confidence is below threshold, or whose sources disagree, is not published at all
|
||||||
approves it.
|
until a human approves it.
|
||||||
- Every event links to its source so a skeptical user can verify in one click.
|
- Every event links to its source so a skeptical user can verify in one click.
|
||||||
|
|
||||||
An empty calendar is a recoverable disappointment. A confidently wrong end date is the failure this
|
An empty calendar is a recoverable disappointment. A confidently wrong end date is the failure this
|
||||||
@@ -120,4 +122,4 @@ product exists to prevent.
|
|||||||
- Should events the user has hidden by game filter still count toward "ends soonest"?
|
- Should events the user has hidden by game filter still count toward "ends soonest"?
|
||||||
(Assumption: no — the filter is global.)
|
(Assumption: no — the filter is global.)
|
||||||
- Is 6 hours the right refresh cadence? (Assumption: yes; events are announced days ahead, so
|
- Is 6 hours the right refresh cadence? (Assumption: yes; events are announced days ahead, so
|
||||||
sub-hourly refresh buys nothing and costs API spend.)
|
sub-hourly refresh buys nothing and is rude to the sources.)
|
||||||
|
|||||||
Reference in New Issue
Block a user