feat: sanitise every string a source publishes

Scraped text becomes React content, JSON on disk and eventually SQLite
rows, so it gets cleaned at one seam: the parse wrapper in toAdapter().
Every source passes through it, a source added tomorrow is covered
without its author doing anything, and no parser can opt out — parsers
stay pure readers of one site's markup.

Removes script/style/comment content and residual tags, decodes entities
to a fixed point so an encoded tag cannot resurrect in a later decoder,
NFKC-normalises, strips control, zero-width and bidi-override characters
(an RTL override visually spoofs a title), bounds each field to the cap
the schema already declares, and requires sourceUrl to be absolute
http(s).

Three things it will not do: touch a date, drop an event it could clean
instead, or repair in silence — every repair and drop is logged by
default. Event IDs are localStorage keys, so an ID is recomputed only
when a sanitised title actually changed and the ID was minted the
standard way; all seven fixtures pass through unrepaired and
byte-identical, which is the regression guard.

Also stops decodeEntities throwing RangeError on an out-of-range
numeric reference, which would have taken a whole source's events down.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
Lucas Winther
2026-08-15 21:15:13 +02:00
co-authored by Claude Opus 5
parent 2b9338a8b0
commit 18b9652aed
4 changed files with 923 additions and 5 deletions
+15 -1
View File
@@ -1,6 +1,7 @@
import type { GachaEvent, GameId } from "../../shared/schema.ts";
import { mergeEvents, type MergeResult } from "../merge.ts";
import { parserById } from "../parsers/index.ts";
import { sanitizeEvents } from "../sanitize.ts";
import { SIX_HOURS_MS, type Adapter, type ParseContext } from "./types.ts";
/**
@@ -90,7 +91,20 @@ function toAdapter(spec: SourceSpec): Adapter {
`${spec.id}: document does not match the '${parser.label}' template; the source has likely been redesigned`,
);
}
return parser.parse(html, ctx);
// The trust boundary. Everything a parser produces came from a page we do
// not control, and this is the one place every source passes through:
// `ADAPTERS` is built from `SOURCES` via this function, so a source added
// tomorrow is sanitised without its author doing anything, and a parser
// cannot opt out. Sanitising here rather than inside the parsers also
// keeps parsers what they are — pure readers of one site's markup.
//
// `sanitizeEvents` logs to console.warn by default, so a repaired or
// dropped event is never silent (CLAUDE.md § Silent drops).
return sanitizeEvents(parser.parse(html, ctx), {
sourceId: ctx.sourceId,
fallbackUrl: ctx.sourceUrl,
}).events;
},
};
}