Add the documentation and the agent instructions
README for humans, AGENTS.md for agents and contributors, and the docs that outlive any one platform: architecture, secrets, connectivity, triggering a deploy from another repository. CLAUDE.md, GEMINI.md and .claude/skills/ are pointers rather than copies, so every agent and every human reads the same text. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 5
parent
971034e630
commit
e1f152ec3e
@@ -0,0 +1,256 @@
|
||||
# compose-stack-deploy
|
||||
|
||||
A small, complete pattern for deploying Docker Compose stacks to a handful of
|
||||
Linux hosts with Ansible — and for letting CI do it safely, one stack at a time,
|
||||
on purpose rather than on every push.
|
||||
|
||||
It is deliberately not a framework. There is **one** Ansible role, about 90 lines
|
||||
of tasks, and a convention about where files live. Everything else in this
|
||||
repository is documentation and six worked examples.
|
||||
|
||||
```
|
||||
server/app/webapp/ → rsync → /srv/stacks/webapp/ on app
|
||||
ansible/playbooks/webapp.yml docker compose up -d
|
||||
```
|
||||
|
||||
## Is this for you?
|
||||
|
||||
It fits if you have somewhere between two and twenty services on a few hosts,
|
||||
each one a `docker-compose.yml`, and you are currently deploying them by hand or
|
||||
with a shell script you no longer trust.
|
||||
|
||||
It is the wrong tool if you want orchestration, scheduling, or rolling updates.
|
||||
That is Kubernetes or Nomad, and you should use one of those.
|
||||
|
||||
## The idea, in three parts
|
||||
|
||||
**Payloads are data.** A stack's compose file and its configuration live under
|
||||
`server/<host>/<stack>/` and are copied to the host *verbatim*. Nothing is
|
||||
templated. You can read the file in the repo and know exactly what is on the
|
||||
host, and you can `docker compose up` it by hand in an emergency.
|
||||
|
||||
**One role, one contract.** Everything a stack can need — networks to create,
|
||||
directories to pre-create, which compose files, what to exclude from the sync,
|
||||
whether to pull — is a variable on a single role. A new service is usually four
|
||||
lines:
|
||||
|
||||
```yaml
|
||||
- name: Deploy the static site
|
||||
hosts: app
|
||||
gather_facts: false
|
||||
tags: [static-site]
|
||||
|
||||
roles:
|
||||
- role: compose_stack
|
||||
vars:
|
||||
stack_name: static-site
|
||||
stack_networks: [proxy-net]
|
||||
```
|
||||
|
||||
**One key, one seam.** Ansible knows exactly one thing about credentials:
|
||||
`SSH_KEY_PATH` points at a `0600` private key. A single script decides where that
|
||||
key comes from — a file, a CI secret, a vault — so adding a backend never touches
|
||||
a playbook.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
ansible/
|
||||
├── ansible.cfg # read only from this directory — see below
|
||||
├── requirements.yml # Galaxy collections
|
||||
├── inventory/
|
||||
│ ├── hosts.yml
|
||||
│ └── group_vars/all.yml # connection settings, stack_root, the pull default
|
||||
├── roles/compose_stack/ # the one shared "sync payload, bring stack up" role
|
||||
└── playbooks/
|
||||
├── site.yml # every stack, in dependency order
|
||||
└── <stack>.yml # one per stack, runnable on its own
|
||||
server/<host>/<stack>/ # the payload: compose files and config, copied verbatim
|
||||
server/shared/ # payloads used by more than one host
|
||||
scripts/load-ssh-key.sh # produces the deploy key path (file / CI / vault)
|
||||
scripts/check.sh # the full static gate — what CI runs
|
||||
skills/adding-a-stack/ # step-by-step guide to adding a service
|
||||
docs/ # secrets, CI setup, connectivity, triggering deploys
|
||||
AGENTS.md # the conventions and the role contract, in full
|
||||
```
|
||||
|
||||
## Getting started
|
||||
|
||||
> **Run every ansible command from the `ansible/` directory.** `ansible.cfg` is
|
||||
> only discovered in the current working directory, and it is what points at the
|
||||
> inventory and puts `roles/` on the roles path. From the repository root you
|
||||
> will get `The role 'compose_stack' was not found`.
|
||||
>
|
||||
> Alternatively export `ANSIBLE_CONFIG=/path/to/repo/ansible/ansible.cfg`, which
|
||||
> works from anywhere. CI has to do this: Ansible ignores an `ansible.cfg` that
|
||||
> lives in a world-writable directory, and a runner's build directory is exactly
|
||||
> that.
|
||||
|
||||
```bash
|
||||
cd ansible
|
||||
ansible-galaxy install -r requirements.yml
|
||||
export SSH_KEY_PATH=~/.ssh/your_deploy_key # see Secrets, below
|
||||
```
|
||||
|
||||
Then check what a deploy *would* do before doing it:
|
||||
|
||||
```bash
|
||||
ansible-playbook --check --diff playbooks/metrics.yml
|
||||
```
|
||||
|
||||
### On Windows
|
||||
|
||||
Ansible has no native Windows control node — `ansible-core` needs a POSIX system,
|
||||
and the scripts here are shell scripts. Use **WSL2** and run everything inside
|
||||
it; the commands above then work unchanged.
|
||||
|
||||
Two things that will cost you an afternoon if you skip them:
|
||||
|
||||
- **Clone into the WSL filesystem** (`~/config`), not `/mnt/c/...`. On the
|
||||
Windows drive `chmod 600` does not stick, and OpenSSH refuses a private key
|
||||
whose permissions are too open. It is also markedly slower.
|
||||
- **Never open a private key or a shell script in Notepad.** It rewrites the file
|
||||
with CRLF line endings, which makes SSH reject the key and makes deployed
|
||||
scripts fail on the host with `bad interpreter: /bin/bash^M`. The
|
||||
`.gitattributes` here keeps LF endings on checkout, but it cannot protect a
|
||||
file you edit outside git.
|
||||
|
||||
## Deploying
|
||||
|
||||
```bash
|
||||
# One stack
|
||||
ansible-playbook playbooks/metrics.yml
|
||||
|
||||
# Everything, in dependency order
|
||||
ansible-playbook playbooks/site.yml
|
||||
|
||||
# One stack out of site.yml
|
||||
ansible-playbook playbooks/site.yml --tags metrics
|
||||
```
|
||||
|
||||
### Deploy versus update
|
||||
|
||||
This distinction is the one piece of the design worth internalising.
|
||||
|
||||
A plain run brings the stack up with the images already on the host. It never
|
||||
pulls, migrates, backs up, or destroys anything. Run it as often as you like.
|
||||
|
||||
```bash
|
||||
ansible-playbook playbooks/webapp.yml -e pull=always
|
||||
```
|
||||
|
||||
`pull=always` fetches newer images and recreates changed containers. It is *also*
|
||||
the switch for the heavier per-stack work: `webapp` only runs its backup and
|
||||
database migration on an update run.
|
||||
|
||||
If your images are tagged `latest`, a plain deploy will **not** pick up a new
|
||||
build — Compose is content with the image it already has. Use `pull=always` when
|
||||
you mean to update.
|
||||
|
||||
Verify with `--check --diff` first. These are production servers.
|
||||
|
||||
## The examples
|
||||
|
||||
Six stacks, each chosen to demonstrate one thing. Read them in this order.
|
||||
|
||||
| Stack | Hosts | What it shows |
|
||||
| --- | --- | --- |
|
||||
| [`static-site`](ansible/playbooks/static-site.yml) | app | The minimum: a name and a network. **Copy this one.** |
|
||||
| [`reverse-proxy`](ansible/playbooks/reverse-proxy.yml) | edge | Pre-creating bind mounts; reloading after a config-only change |
|
||||
| [`metrics`](ansible/playbooks/metrics.yml) | edge, app | One payload shared by two hosts; per-host variable maps; directories owned by a container's uid |
|
||||
| [`webapp`](ansible/playbooks/webapp.yml) | app | Backup and migrate *between* sync and up, gated on an update run; a second compose file; building from a Dockerfile; requiring a host-side `.env` |
|
||||
| [`webapp-staging`](ansible/playbooks/webapp-staging.yml) | app | A second instance of the same payload, elsewhere, with different settings |
|
||||
| [`banner`](ansible/playbooks/banner.yml) | edge, app | The escape hatch: something that is not a Compose stack at all |
|
||||
|
||||
`server/` is data — see [`server/README.md`](server/README.md).
|
||||
|
||||
## Secrets
|
||||
|
||||
Three separate things, handled three different ways. Full detail in
|
||||
[`docs/secrets.md`](docs/secrets.md).
|
||||
|
||||
**1. The deploy SSH key.** Ansible only ever reads `SSH_KEY_PATH`.
|
||||
`scripts/load-ssh-key.sh` is the only thing that knows where the key comes from;
|
||||
it prints a path and never prints key material.
|
||||
|
||||
```bash
|
||||
export SSH_KEY_PATH="$(scripts/load-ssh-key.sh)" # key already on disk
|
||||
export SSH_KEY_PATH="$(scripts/load-ssh-key.sh ci)" # from a CI secret
|
||||
export SSH_KEY_PATH="$(scripts/load-ssh-key.sh vault)" # from Vault / OpenBao
|
||||
```
|
||||
|
||||
**2. Per-stack `.env` files.** Stack secrets live on the hosts, not here. The
|
||||
role excludes `.env` from every sync, so editing one in the repo has no effect
|
||||
and a host's `.env` is never overwritten. The example stacks use Compose's
|
||||
`${VAR:?message}` syntax, so a stack refuses to start with a missing credential
|
||||
rather than booting with an empty password.
|
||||
|
||||
**3. Nothing else.** There is no third category, and that is the point. A secret
|
||||
written directly into a compose file cannot be excluded from the sync, cannot be
|
||||
rotated without a commit, and is in every clone of the repository forever. CI
|
||||
runs [gitleaks](https://github.com/gitleaks/gitleaks) over the working tree *and*
|
||||
the history on every push to keep it that way.
|
||||
|
||||
## CI/CD
|
||||
|
||||
All three platforms are maintained, and all three do the same two things.
|
||||
|
||||
| | GitHub Actions | GitLab CI | Gitea Actions |
|
||||
| --- | --- | --- | --- |
|
||||
| Static gate | [`lint.yml`](.github/workflows/lint.yml) | `lint` job | [`lint.yml`](.gitea/workflows/lint.yml) |
|
||||
| Deploy | [`deploy.yml`](.github/workflows/deploy.yml) | per-stack buttons | [`deploy.yml`](.gitea/workflows/deploy.yml) |
|
||||
| Diagnostics | [`connectivity.yml`](.github/workflows/connectivity.yml) | `ansible:ping` | [`connectivity.yml`](.gitea/workflows/connectivity.yml) |
|
||||
| Setup guide | [`ci-github-actions.md`](docs/ci-github-actions.md) | [`ci-gitlab.md`](docs/ci-gitlab.md) | [`ci-gitea.md`](docs/ci-gitea.md) |
|
||||
|
||||
Gitea reads `.gitea/workflows/` *instead of* `.github/workflows/` — it stops at
|
||||
the first workflow directory that exists — so the two never both run. Delete the
|
||||
platforms you are not using rather than letting them rot.
|
||||
|
||||
The rules that matter, on all three:
|
||||
|
||||
- **The lint gate runs on every push** and must pass. It needs no access to any
|
||||
host. It is `scripts/check.sh` — the same command you run locally — so it
|
||||
cannot drift from what you tested.
|
||||
- **Deploys are never automatic.** Merging changes what *would* be deployed; a
|
||||
person still decides when. Every deploy path is manual or externally triggered.
|
||||
- **Check before deploy.** All three expose a check mode that runs
|
||||
`--check --diff` and touches nothing.
|
||||
- **To update rather than deploy**, set the pull policy to `always`.
|
||||
- **Untrusted input is validated.** The stack name and pull policy reach a shell
|
||||
command and, for externally triggered deploys, come from another repository.
|
||||
Both are matched against a fixed pattern before anything runs.
|
||||
|
||||
Another repository can ask this one to redeploy a stack after building a new
|
||||
image — see [`docs/triggering-deploys.md`](docs/triggering-deploys.md).
|
||||
|
||||
## Adding a stack
|
||||
|
||||
1. Put the compose file and its configuration in `server/<host>/<stack>/`.
|
||||
2. Add `ansible/playbooks/<stack>.yml`. Copy `static-site.yml`.
|
||||
3. Add it to `site.yml`, in dependency order.
|
||||
4. Add it to CI: the `stack` dropdown in `.github/workflows/deploy.yml` and in
|
||||
`.gitea/workflows/deploy.yml`, and a `check:`/`deploy:` pair in
|
||||
`.gitlab-ci.yml`. Skip this and the stack works from a workstation but is
|
||||
invisible in CI, which is how services end up hand-deployed.
|
||||
5. `make check`, then `--check --diff` against the host.
|
||||
|
||||
The long version, with a table of which contract variable solves which problem,
|
||||
is in [`skills/adding-a-stack/SKILL.md`](skills/adding-a-stack/SKILL.md).
|
||||
[`AGENTS.md`](AGENTS.md) has the conventions and the full role contract.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Symptom | Cause |
|
||||
| --- | --- |
|
||||
| `The role 'compose_stack' was not found` | Not running from the `ansible/` directory |
|
||||
| `SSH_KEY_PATH does not point at a readable file` | Key not loaded, or a CI job that cannot see the secret |
|
||||
| A deploy succeeds but the service is unchanged | Image tag is `latest` and you did not pass `-e pull=always` |
|
||||
| Config file copied but nothing happens | Bind-mounted config; the stack needs a reload, which the playbook does only when the sync changed something |
|
||||
| `no payload directory at ...` | `stack_src` does not match where the files actually are |
|
||||
| A stack refuses to start, naming a variable | Its `.env` is missing on the host — that is the design working |
|
||||
| `Permission denied (publickey)` | The matching public key is not in the deploy user's `authorized_keys` |
|
||||
| `Connection timed out` | A network path problem, not a key problem — see [`docs/connectivity.md`](docs/connectivity.md) |
|
||||
|
||||
## License
|
||||
|
||||
MIT. See [LICENSE](LICENSE).
|
||||
Reference in New Issue
Block a user