fix(refresh): make a source that stopped answering turn the run red
The only snapshot the scheduled refresh has ever committed is endfield-wikigg-events. Seven sources existed at that run; the six game8.co ones yielded nothing, and no cycle since has committed anything. All of them fetch fine from a laptop, so whatever is happening happens on the runner — meanwhile eight of ten games were served from checked-in fixtures for three days behind a green tick. Nine failures out of ten was exit 0 with warnings buried in a log nobody opens. `consecutiveFailures` was already tracked and never read. A source that has failed BROKEN_AFTER_FAILURES (3, so ~36h at two cycles a day) is now reported as `broken`: a GitHub annotation, a job-summary row carrying its status code, and a `broken` step output. The runner still exits 0 on it and `refresh.yml` fails on that output in a final step, after the commit and the CI dispatch — exiting non-zero from the runner would skip the commit and throw away the pages that did arrive, which is the opposite of what "one wiki down never blanks a calendar" is for. The streak is read from the store rather than from this cycle's outcome, so a source dead for days that happens to be inside its six-hour window has not recovered. A non-ok response now records what turned us away — the Server header, whether a CF-Ray was present, any Retry-After — because a bare `HTTP 403` reads identically whether the page moved behind a login or a CDN decided the runner is a bot farm, and that is the open question here. Values are trimmed and capped: the note lands in a workflow command and a markdown cell, and it came from a host we do not control. Also space requests to a host already asked this cycle, honouring its Crawl-delay and defaulting to 2s. Eight sources share game8.co, so the per-source floor alone still permitted one cycle to arrive as eight back-to-back requests to a single site — which is what a burst looks like from the far end regardless of our intent, and is plausibly self-inflicted here. The wait is taken after the interval and robots gates, so a source we then skip costs nothing. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 5
parent
d2615c606c
commit
2a0ea1a796
@@ -71,6 +71,7 @@ jobs:
|
||||
restore-keys: refresh-state-
|
||||
|
||||
- name: Refresh
|
||||
id: refresh
|
||||
env:
|
||||
# Identify the crawler with a contact URL, per CLAUDE.md.
|
||||
REFRESH_CONTACT_URL: ${{ github.server_url }}/${{ github.repository }}
|
||||
@@ -155,3 +156,28 @@ jobs:
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: gh workflow run ci.yml --ref "${{ github.ref_name }}"
|
||||
|
||||
# A source that has failed three cycles running is broken, not down: that
|
||||
# game's calendar has been built from a checked-in fixture for a day and a
|
||||
# half while every run showed a green tick. `bun run refresh` exits 0 on
|
||||
# this so the steps above still commit and publish what did work; turning
|
||||
# the run red is this step's job, and it is last for that reason.
|
||||
#
|
||||
# `always()` so it still reports when an earlier step failed — but note it
|
||||
# cannot report when the Refresh step itself hard-failed, since the output
|
||||
# is then unset and the job is already red on its own account.
|
||||
- name: Report source health
|
||||
if: always()
|
||||
env:
|
||||
BROKEN: ${{ steps.refresh.outputs.broken }}
|
||||
REFRESH_OUTCOME: ${{ steps.refresh.outcome }}
|
||||
run: |
|
||||
if [ "$REFRESH_OUTCOME" != "success" ]; then
|
||||
echo "refresh did not complete ($REFRESH_OUTCOME); no health to report"
|
||||
exit 0
|
||||
fi
|
||||
if [ -n "$BROKEN" ] && [ "$BROKEN" != "0" ]; then
|
||||
echo "::error::$BROKEN source(s) have stopped answering; see the job summary"
|
||||
exit 1
|
||||
fi
|
||||
echo "every source is answering"
|
||||
|
||||
Reference in New Issue
Block a user