CI/CD · GitOps

ArgoCD + GitHub Actions: the pipeline we ship by default

TL;DR

  • GitHub Actions builds. ArgoCD deploys. One pipeline doing both is where most bad nights start.
  • The only thing handed between them is an image tag written into Git. No cluster credentials in CI.
  • Self-heal undoes anything changed by hand, so the repository stays a truthful record of production.
  • Rollback is a revert, not a rebuild. Seconds, not another full pipeline run.
ArgoCD + GitHub Actions: the pipeline we ship by default — CloudDrove

The pipeline that pages you at 3am

Almost every team we inherit has one workflow file that does everything. It tests, it builds, it logs into the cluster, it runs kubectl apply, and on a good day it works.

Then a Friday comes along. A deploy half succeeded. Someone patched a deployment by hand to stop the bleeding, and forgot. The workflow that "always worked" now applies an old manifest over the fix. Nobody can answer the only question that matters: what is actually running in production right now?

The repository says one thing. The cluster says another. And the CI system, which is the least protected part of the stack, is holding admin credentials for both.

Split the job in two

The fix is boring, and that is the point. Two jobs, two owners, one contract between them.

Pipeline flow: GitHub Actions tests and builds, pushes the image to the registry, writes the tag into the config repo, ArgoCD applies it to the cluster. A healthy sync goes live, an unhealthy one is reverted in Git.
CI never touches the cluster. It writes a tag, the agent inside the cluster does the rest, and a bad sync goes back through Git.

GitHub Actions owns everything up to the image. Tests, build, scan, push. It is allowed to fail loudly, because nothing it does is visible to users yet.

ArgoCD owns everything after. It runs inside the cluster, watches a Git repository, and makes the cluster match it. If the two disagree, Git wins.

That single rule removes the whole category of "who changed what". It also removes cluster credentials from CI, which is the change your security reviewer will care about most.

Step 1: build once, tag by commit

Build the image once and tag it with the commit SHA. Never latest. A tag that can point at two different images is a tag you cannot roll back to.

# .github/workflows/build.yml
name: build
on:
  push:
    branches: [main]

jobs:
  image:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      id-token: write          # OIDC, no long-lived AWS keys
    steps:
      - uses: actions/checkout@v4
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: ${{ secrets.ECR_PUSH_ROLE }}
          aws-region: ca-central-1
      - uses: aws-actions/amazon-ecr-login@v2
      - run: |
          docker build -t $REPO:${{ github.sha }} .
          docker push $REPO:${{ github.sha }}

Two details worth copying. The job authenticates to AWS with OIDC, so there are no permanent access keys sitting in GitHub. And the image is pushed to your own registry, not pulled from Docker Hub at deploy time, which is how you avoid the rate limit outage we wrote about here.

Step 2: CI writes one line

Now the contract. The build job's last act is to write the new tag into the configuration repository. One line, one commit, nothing else.

# still in the same workflow, after the push
      - run: |
          git clone https://x:${{ secrets.CONFIG_REPO_TOKEN }}@github.com/acme/deploy.git
          cd deploy/apps/checkout/overlays/prod
          kustomize edit set image app=$REPO:${{ github.sha }}
          git commit -am "checkout: ${{ github.sha }}"
          git push

Keep application code and deployment configuration in two repositories. It feels like extra work for about a week. After that, it means the history of what shipped to production is a readable list of commits, separate from the noise of feature branches.

Step 3: ArgoCD notices

ArgoCD watches that repository. When the tag changes, it applies the change and reports health. The application definition is small, and the three lines that matter are at the bottom.

apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: checkout
  namespace: argocd
spec:
  project: default
  source:
    repoURL: https://github.com/acme/deploy.git
    targetRevision: main
    path: apps/checkout/overlays/prod
  destination:
    server: https://kubernetes.default.svc
    namespace: checkout
  syncPolicy:
    automated:
      prune: true         # delete what Git no longer declares
      selfHeal: true      # undo manual kubectl changes
    retry:
      limit: 3

selfHeal is the line that ends the Friday night story. Patch a deployment by hand and ArgoCD puts it back within seconds. If the hand patch was right, it belongs in Git, and now there is pressure to put it there.

prune is the one teams enable last, and regret waiting on. Without it, resources you delete from Git keep running in the cluster forever, quietly costing money.

What actually goes wrong

Four failures, in the order we see them.

1. Secrets in the config repo. The moment deployment config lives in Git, someone commits a password. Use External Secrets or Sealed Secrets so the repository holds a reference, never a value. Decide this on day one, not after a leak.

2. Health checks that always pass. ArgoCD reports "Healthy" based on what the resources say about themselves. A deployment with no readiness probe is healthy the second the pod starts, even while it returns errors. Probes are part of the pipeline, not a nice-to-have.

3. One giant application. Thirty services in one ArgoCD Application means one bad manifest blocks every deploy. Split by service, and let each team break only their own thing.

4. Nobody reads the sync failures. A failed sync is a silent stop, not an outage, so it sits. Route ArgoCD notifications into the same channel as your alerts, and treat a failed sync like a failed build. On our own pipelines Naoru reads the failed run and posts the likely cause on the pull request, so the fix starts before anyone opens the logs.

Rollback is a revert

The reward for all of this arrives the first time something ships broken. You do not rebuild. You do not hunt for the last good artifact. You revert the commit that changed the tag.

git revert HEAD && git push   # ArgoCD syncs back to the previous image

The old image is still in the registry, so recovery takes the length of a sync rather than the length of a pipeline. That is usually under a minute, and it works the same at 3am as it does at 3pm, which is the entire point.

What to do next

01

Take cluster credentials out of CI first. Even before you install anything, that single change removes the worst blast radius. Our CI/CD & GitOps work usually starts here.

02

Pick your agent on how your team works, not on benchmarks. We compared both honestly in ArgoCD vs Flux.

03

Decide how traffic moves once the sync happens. Rolling is fine for most services; for the rest, see blue/green on EKS.

All blogs

Related Reading

Go deeper.

Cloud Infrastructure Assessment

See exactly where your cloud stands.

A senior engineer reviews your architecture, cost, security, and reliability, then sends back a prioritized findings report, the fixes that matter most, in order.

  • Architecture & scale
  • Cost & efficiency
  • Security & reliability
Book an Assessment

Complimentary · no obligation · no sales pressure

Work With Us

Want this kind of engineering on your side?

The same people who write these build your platform. Let's talk about what you're working on.

Talk to an Expert