feat: sanitise every string a source publishes
Scraped text becomes React content, JSON on disk and eventually SQLite rows, so it gets cleaned at one seam: the parse wrapper in toAdapter(). Every source passes through it, a source added tomorrow is covered without its author doing anything, and no parser can opt out — parsers stay pure readers of one site's markup. Removes script/style/comment content and residual tags, decodes entities to a fixed point so an encoded tag cannot resurrect in a later decoder, NFKC-normalises, strips control, zero-width and bidi-override characters (an RTL override visually spoofs a title), bounds each field to the cap the schema already declares, and requires sourceUrl to be absolute http(s). Three things it will not do: touch a date, drop an event it could clean instead, or repair in silence — every repair and drop is logged by default. Event IDs are localStorage keys, so an ID is recomputed only when a sanitised title actually changed and the ID was minted the standard way; all seven fixtures pass through unrepaired and byte-identical, which is the regression guard. Also stops decodeEntities throwing RangeError on an out-of-range numeric reference, which would have taken a whole source's events down. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 5
parent
2b9338a8b0
commit
18b9652aed
@@ -1,6 +1,7 @@
|
||||
import type { GachaEvent, GameId } from "../../shared/schema.ts";
|
||||
import { mergeEvents, type MergeResult } from "../merge.ts";
|
||||
import { parserById } from "../parsers/index.ts";
|
||||
import { sanitizeEvents } from "../sanitize.ts";
|
||||
import { SIX_HOURS_MS, type Adapter, type ParseContext } from "./types.ts";
|
||||
|
||||
/**
|
||||
@@ -90,7 +91,20 @@ function toAdapter(spec: SourceSpec): Adapter {
|
||||
`${spec.id}: document does not match the '${parser.label}' template; the source has likely been redesigned`,
|
||||
);
|
||||
}
|
||||
return parser.parse(html, ctx);
|
||||
|
||||
// The trust boundary. Everything a parser produces came from a page we do
|
||||
// not control, and this is the one place every source passes through:
|
||||
// `ADAPTERS` is built from `SOURCES` via this function, so a source added
|
||||
// tomorrow is sanitised without its author doing anything, and a parser
|
||||
// cannot opt out. Sanitising here rather than inside the parsers also
|
||||
// keeps parsers what they are — pure readers of one site's markup.
|
||||
//
|
||||
// `sanitizeEvents` logs to console.warn by default, so a repaired or
|
||||
// dropped event is never silent (CLAUDE.md § Silent drops).
|
||||
return sanitizeEvents(parser.parse(html, ctx), {
|
||||
sourceId: ctx.sourceId,
|
||||
fallbackUrl: ctx.sourceUrl,
|
||||
}).events;
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user