CI/CD · GitOps
ArgoCD + GitHub Actions: the pipeline we ship by default
TL;DR
- GitHub Actions builds. ArgoCD deploys. One pipeline doing both is where most bad nights start.
- The only thing handed between them is an image tag written into Git. No cluster credentials in CI.
- Self-heal undoes anything changed by hand, so the repository stays a truthful record of production.
- Rollback is a revert, not a rebuild. Seconds, not another full pipeline run.
The pipeline that pages you at 3am
Almost every team we inherit has one workflow file that does everything. It tests, it builds, it logs into the cluster, it runs kubectl apply, and on a good day it works.
Then a Friday comes along. A deploy half succeeded. Someone patched a deployment by hand to stop the bleeding, and forgot. The workflow that "always worked" now applies an old manifest over the fix. Nobody can answer the only question that matters: what is actually running in production right now?
The repository says one thing. The cluster says another. And the CI system, which is the least protected part of the stack, is holding admin credentials for both.
Split the job in two
The fix is boring, and that is the point. Two jobs, two owners, one contract between them.
GitHub Actions owns everything up to the image. Tests, build, scan, push. It is allowed to fail loudly, because nothing it does is visible to users yet.
ArgoCD owns everything after. It runs inside the cluster, watches a Git repository, and makes the cluster match it. If the two disagree, Git wins.
That single rule removes the whole category of "who changed what". It also removes cluster credentials from CI, which is the change your security reviewer will care about most.
Step 1: build once, tag by commit
Build the image once and tag it with the commit SHA. Never latest. A tag that can point at two different images is a tag you cannot roll back to.
# .github/workflows/build.yml
name: build
on:
push:
branches: [main]
jobs:
image:
runs-on: ubuntu-latest
permissions:
contents: read
id-token: write # OIDC, no long-lived AWS keys
steps:
- uses: actions/checkout@v4
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: ${{ secrets.ECR_PUSH_ROLE }}
aws-region: ca-central-1
- uses: aws-actions/amazon-ecr-login@v2
- run: |
docker build -t $REPO:${{ github.sha }} .
docker push $REPO:${{ github.sha }}Two details worth copying. The job authenticates to AWS with OIDC, so there are no permanent access keys sitting in GitHub. And the image is pushed to your own registry, not pulled from Docker Hub at deploy time, which is how you avoid the rate limit outage we wrote about here.
Step 2: CI writes one line
Now the contract. The build job's last act is to write the new tag into the configuration repository. One line, one commit, nothing else.
# still in the same workflow, after the push
- run: |
git clone https://x:${{ secrets.CONFIG_REPO_TOKEN }}@github.com/acme/deploy.git
cd deploy/apps/checkout/overlays/prod
kustomize edit set image app=$REPO:${{ github.sha }}
git commit -am "checkout: ${{ github.sha }}"
git pushKeep application code and deployment configuration in two repositories. It feels like extra work for about a week. After that, it means the history of what shipped to production is a readable list of commits, separate from the noise of feature branches.
Step 3: ArgoCD notices
ArgoCD watches that repository. When the tag changes, it applies the change and reports health. The application definition is small, and the three lines that matter are at the bottom.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: checkout
namespace: argocd
spec:
project: default
source:
repoURL: https://github.com/acme/deploy.git
targetRevision: main
path: apps/checkout/overlays/prod
destination:
server: https://kubernetes.default.svc
namespace: checkout
syncPolicy:
automated:
prune: true # delete what Git no longer declares
selfHeal: true # undo manual kubectl changes
retry:
limit: 3selfHeal is the line that ends the Friday night story. Patch a deployment by hand and ArgoCD puts it back within seconds. If the hand patch was right, it belongs in Git, and now there is pressure to put it there.
prune is the one teams enable last, and regret waiting on. Without it, resources you delete from Git keep running in the cluster forever, quietly costing money.
What actually goes wrong
Four failures, in the order we see them.
1. Secrets in the config repo. The moment deployment config lives in Git, someone commits a password. Use External Secrets or Sealed Secrets so the repository holds a reference, never a value. Decide this on day one, not after a leak.
2. Health checks that always pass. ArgoCD reports "Healthy" based on what the resources say about themselves. A deployment with no readiness probe is healthy the second the pod starts, even while it returns errors. Probes are part of the pipeline, not a nice-to-have.
3. One giant application. Thirty services in one ArgoCD Application means one bad manifest blocks every deploy. Split by service, and let each team break only their own thing.
4. Nobody reads the sync failures. A failed sync is a silent stop, not an outage, so it sits. Route ArgoCD notifications into the same channel as your alerts, and treat a failed sync like a failed build. On our own pipelines Naoru reads the failed run and posts the likely cause on the pull request, so the fix starts before anyone opens the logs.
Rollback is a revert
The reward for all of this arrives the first time something ships broken. You do not rebuild. You do not hunt for the last good artifact. You revert the commit that changed the tag.
git revert HEAD && git push # ArgoCD syncs back to the previous imageThe old image is still in the registry, so recovery takes the length of a sync rather than the length of a pipeline. That is usually under a minute, and it works the same at 3am as it does at 3pm, which is the entire point.
