`lastConfirmedAt` lives in gitignored bookkeeping that only refresh.yml restored, so the workflow that actually builds and deploys never saw it: every source reported its last *content change* as its last success, and the UI flagged anything whose bytes had not moved in two days as stale — which is most wiki pages most of the time. ci.yml now restores the same cache read-only before building the feed. The refresh push was a bare `git push`, so a human push landing in between made it non-fast-forward: the job failed and threw away pages it had just fetched, while the bookkeeping had already been saved, so those sources would not be re-asked for six hours. Rebase and retry instead — never force. The cache save key used run_id, which is stable across re-runs, so a re-run saved nothing and the run after it restored stale bookkeeping. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
158 lines
6.3 KiB
YAML
158 lines
6.3 KiB
YAML
name: Refresh sources
|
|
|
|
# Fetch each source at most twice a day and commit the raw snapshots when — and
|
|
# only when — the bytes actually changed. Everything downstream (parse, merge,
|
|
# feed, build, deploy) is CI's existing job; this workflow does not duplicate
|
|
# any of it, it just hands CI fresher input.
|
|
#
|
|
# Twelve hours apart is deliberately well clear of the six-hour-per-source floor
|
|
# in CLAUDE.md § Scraping conduct, and the runner enforces that floor itself, so
|
|
# a manual dispatch on top of a scheduled run cannot double up on a wiki.
|
|
on:
|
|
schedule:
|
|
- cron: "27 5,17 * * *"
|
|
workflow_dispatch:
|
|
inputs:
|
|
dry_run:
|
|
description: "Plan only — no requests, no writes"
|
|
type: boolean
|
|
default: false
|
|
only:
|
|
description: "Refresh a single source id (blank = all)"
|
|
type: string
|
|
default: ""
|
|
|
|
# Never two refreshes at once: they would both fetch, and the second would race
|
|
# the first's commit. Queue instead of cancelling — a half-finished refresh that
|
|
# has already written snapshots should be allowed to finish and push.
|
|
concurrency:
|
|
group: refresh
|
|
cancel-in-progress: false
|
|
|
|
permissions:
|
|
contents: read
|
|
|
|
jobs:
|
|
refresh:
|
|
name: Fetch sources and rebuild the feed
|
|
runs-on: ubuntu-latest
|
|
# A fork must not point this at the wikis on a schedule. Same shape as the
|
|
# DEPLOY_PAGES gate in ci.yml: off by default for anyone but this repo,
|
|
# while a fork owner can still dispatch it by hand and take responsibility.
|
|
if: >-
|
|
github.event_name == 'workflow_dispatch' ||
|
|
github.repository == 'StereotypicalCat/gacha-event-tracker'
|
|
permissions:
|
|
# Commit the refreshed snapshots.
|
|
contents: write
|
|
# Dispatch ci.yml afterwards: a push made with GITHUB_TOKEN deliberately
|
|
# does not trigger other workflows, so without this the fresh data would
|
|
# sit in the repo undeployed until someone pushed by hand.
|
|
actions: write
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
|
|
- uses: oven-sh/setup-bun@v2
|
|
with:
|
|
bun-version: "1.3"
|
|
|
|
- run: bun install --frozen-lockfile
|
|
|
|
# When each source was last checked. Gitignored on purpose (committing it
|
|
# would mean a commit every cycle saying nothing changed), so it rides in
|
|
# the actions cache instead. A cache miss only means the runner has no
|
|
# record of the last check — the twelve-hour schedule still keeps us well
|
|
# inside the etiquette floor.
|
|
- name: Restore refresh bookkeeping
|
|
uses: actions/cache/restore@v4
|
|
with:
|
|
path: snapshots/*.state.json
|
|
key: refresh-state-
|
|
restore-keys: refresh-state-
|
|
|
|
- name: Refresh
|
|
env:
|
|
# Identify the crawler with a contact URL, per CLAUDE.md.
|
|
REFRESH_CONTACT_URL: ${{ github.server_url }}/${{ github.repository }}
|
|
# Passed through the environment rather than interpolated into the
|
|
# run script, so a dispatch input cannot become shell.
|
|
ONLY: ${{ inputs.only }}
|
|
DRY_RUN: ${{ inputs.dry_run }}
|
|
run: |
|
|
args=()
|
|
if [ "$DRY_RUN" = "true" ]; then
|
|
args+=(--dry-run)
|
|
fi
|
|
if [ -n "$ONLY" ]; then
|
|
args+=(--only "$ONLY")
|
|
fi
|
|
bun run refresh "${args[@]}"
|
|
|
|
# The key must differ every time this step runs. `run_id` is stable across
|
|
# re-runs, so a re-run's save hits an existing key, is skipped, and the
|
|
# next run restores the bookkeeping from before the re-run — records of
|
|
# requests we did make, lost. `run_attempt` increments per attempt.
|
|
- name: Save refresh bookkeeping
|
|
if: always()
|
|
uses: actions/cache/save@v4
|
|
with:
|
|
path: snapshots/*.state.json
|
|
key: refresh-state-${{ github.run_id }}-${{ github.run_attempt }}
|
|
|
|
# git is the authority on "did anything change" — a 304, an unchanged
|
|
# body, or a rejected parse all leave the working tree clean.
|
|
- name: Detect changes
|
|
id: diff
|
|
run: |
|
|
if [ -n "$(git status --porcelain -- snapshots)" ]; then
|
|
git status --porcelain -- snapshots
|
|
echo "changed=true" >> "$GITHUB_OUTPUT"
|
|
else
|
|
echo "no source changed"
|
|
echo "changed=false" >> "$GITHUB_OUTPUT"
|
|
fi
|
|
|
|
# A human push landing between the checkout and this push makes the push
|
|
# non-fast-forward. Failing there would throw away pages we have already
|
|
# fetched while the bookkeeping above (saved with `if: always()`) has
|
|
# already spent their six-hour budget — the wikis would be asked again for
|
|
# nothing. So rebase onto whatever landed and try again. Never force: this
|
|
# commit is only ever new files under snapshots/, so it has nothing to say
|
|
# about anyone else's work.
|
|
- name: Commit refreshed snapshots
|
|
if: steps.diff.outputs.changed == 'true' && inputs.dry_run != true
|
|
env:
|
|
BRANCH: ${{ github.ref_name }}
|
|
run: |
|
|
set -euo pipefail
|
|
git config user.name "github-actions[bot]"
|
|
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
|
|
git add -- snapshots
|
|
git commit -m "chore(data): refresh source snapshots" \
|
|
-m "Automated fetch from ${{ github.workflow }} run ${{ github.run_id }}."
|
|
|
|
for attempt in 1 2 3; do
|
|
if git push origin "HEAD:$BRANCH"; then
|
|
exit 0
|
|
fi
|
|
echo "push rejected (attempt $attempt); rebasing onto origin/$BRANCH"
|
|
git fetch origin "$BRANCH"
|
|
if ! git rebase "origin/$BRANCH"; then
|
|
git rebase --abort || true
|
|
echo "::error::snapshot commit conflicts with $BRANCH; not force-pushing"
|
|
exit 1
|
|
fi
|
|
sleep $((attempt * 5))
|
|
done
|
|
echo "::error::could not push refreshed snapshots after 3 attempts"
|
|
exit 1
|
|
|
|
# ci.yml owns typecheck, tests, the feed sanity check, the image and the
|
|
# Pages deploy. Dispatching it is how the refreshed data reaches the site
|
|
# without any of that logic being copied here.
|
|
- name: Publish the refreshed feed
|
|
if: steps.diff.outputs.changed == 'true' && inputs.dry_run != true
|
|
env:
|
|
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
|
run: gh workflow run ci.yml --ref "${{ github.ref_name }}"
|