Summary
How do you achieve 100% uptime when code changes daily? This article covers the architecture required to eliminate downtime in business-critical systems. The client in this case study preferred to remain anonymous, and is referred to as "NordFinans".
The focus is on the technical implementation of Zero Downtime Deployment (ZDD): how to configure Kubernetes, database schemas, and application logic for continuous delivery without interruption, whether your stack is built on PHP, Python, Go, or Node.js.
Deployment frequency
Lead time (commit → prod)
Background: The Fear of Deployment
The Old World
Many companies will recognize the situation NordFinans faced: a large monolithic application that had grown unwieldy, and a release process that was manual, slow, and risky:
Maintenance windows: Updates required planned downtime, typically scheduled for late evenings. This created frustration for users who expected 24/7 services, and wore out developers who had to work nights.
"Deployment Fear": Because each rollout was a major operation, they were postponed as long as possible. This led to a vicious cycle of enormous code conflicts and increased probability of errors.
Inefficient scaling: During traffic peaks, the entire monolith had to be scaled up, even if only a small part of the system was under pressure.
The Goal
To meet today's availability requirements, the company set three absolute technical requirements:
| Requirement | Goal |
|---|---|
| Deployment Frequency | From monthly to daily rollouts |
| Change Failure Rate | Under 1% failure rate during deployment |
| True Zero Downtime | No interrupted sessions or 5xx errors |
Strategy: Hybrid CI/CD with GitLab and GitHub
To balance internal security with external availability, NordFinans went with a hybrid strategy. The model suits companies that have both a proprietary core business and public integrations.
GitLab: The Core for DevSecOps
GitLab (Self-Managed) is the primary platform for internal source code and infrastructure.
Why: GitLab offers a complete package with source code, CI pipelines, container registry, and security scanning (SAST/DAST) in one closed ecosystem.
Kubernetes integration: Via GitLab Agent, the platform team can control access granularly without exposing sensitive access keys to developers.
GitHub: The Public Face
Public SDKs and partner integrations live on GitHub.
Why: GitHub is the industry standard for open-source.
GitHub Actions: Actions run the public tests and publish packages to registries like NPM, PyPI, and Packagist.
Synchronization
To avoid fragmentation, GitLab is the "Source of Truth". Code is mirrored automatically to GitHub, so developers only deal with one dashboard while the code lives in two places.
CI/CD Architecture
GitOps: The Engine Under the Hood
Zero downtime starts with removing the manual sources of error. Nobody runs kubectl apply by hand anymore; everything goes through a pure GitOps model.
ArgoCD as Traffic Police
ArgoCD keeps the state in the Kubernetes cluster in sync with the state in Git.
Pull-based model: Instead of the CI server "pushing" changes to the cluster (which requires the CI server to have admin access to prod), ArgoCD "pulls" changes from a separate manifest repo.
Security: The CI system never has direct access to the production environment, which removes a large attack surface.
The Flow from Code to Prod
- CI (Build): Developer pushes code. Pipeline runs tests, builds Docker image, and scans for vulnerabilities.
- CD (Update): If the build succeeds, the CI job updates the version tag in a separate manifest repo.
- Sync: ArgoCD detects the change, calculates the difference, and rolls out the change in a controlled manner in Kubernetes.
Technical Deep Dive: How to Achieve 100% Uptime?
Replacing the engine on a plane while it's in the air takes precision. These are the configurations that let you roll out new versions during working hours without losing a single request.
The Rolling Update Strategy
The default behavior of Kubernetes is a "Rolling Update", but the default settings are often too aggressive for critical applications. The strategy must be adjusted to guarantee capacity:
apiVersion: apps/v1kind: Deploymentmetadata: name: api-serverspec: replicas: 4 strategy: type: RollingUpdate rollingUpdate: maxSurge: 25% maxUnavailable: 0 template: spec: containers: - name: api image: registry/api:v2.1.0 ports: - containerPort: 8080Graceful Shutdown: The Solution to 502 Bad Gateway
The most common error when transitioning to Kubernetes is ignoring the application's lifecycle. When a pod is about to die, two things happen simultaneously (asynchronously):
- Kubernetes removes the pod's IP from the load balancers.
- Kubernetes sends SIGTERM to the container to stop the process.
The problem: Processes like Nginx, Go binaries, or Node.js often stop faster than Kubernetes can update the network rules across the cluster. Traffic keeps landing on a pod that has just died, and the user sees "502 Bad Gateway".
spec: containers: - name: api lifecycle: preStop: exec: command: ["/bin/sh", "-c", "sleep 15"] # Graceful shutdown in the application terminationGracePeriodSeconds: 30Probes: The Art of Health Checks
Liveness Probe: "Am I alive?". It checks that the process is running, and it should stay simple. Don't check the database connection here: if the database goes down, every pod restarts at the same time in an endless loop.
Readiness Probe: "Am I ready to receive traffic?". This one should check that the application can actually do work (e.g., db connection ok, cache warm). If it fails, the pod is taken out of the traffic flow without being restarted.
spec: containers: - name: api livenessProbe: httpGet: path: /health/live port: 8080 initialDelaySeconds: 10 periodSeconds: 10 readinessProbe: httpGet: path: /health/ready port: 8080 initialDelaySeconds: 5 periodSeconds: 5 failureThreshold: 3The Database: The Biggest Challenge
Code is ephemeral, but data is persistent. How do you update a database schema without locking tables or crashing the old version of the code that's still running during a rollout?
The solution is the Expand-Contract (Parallel Change) pattern.
Phase 1: Expand
Are we changing a column name from address to billing_address? We add the new column but keep the old one. We roll out the code. Now both columns exist.
Phase 2: Migrate (Dual Write)
The application is updated to write to both columns but read from the new one. A background script moves old data.
Phase 3: Contract
When we're sure all pods are running new code that uses billing_address, we remove the old column in a final migration.
// Migration 1: Expand - Add new columnSchema::table('customers', function (Blueprint $table) { $table->string('billing_address')->nullable();}); // Model: Dual write during transition periodclass Customer extends Model{ public function setAddressAttribute($value) { $this->attributes['address'] = $value; $this->attributes['billing_address'] = $value; } public function getAddressAttribute() { return $this->billing_address ?? $this->attributes['address']; }} // Migration 2: Contract - Remove old columnSchema::table('customers', function (Blueprint $table) { $table->dropColumn('address');});This requires discipline, but guarantees that the database is never "out of sync" with any version of the application running live.
Application Level: Preparing for the Cloud
Regardless of whether the application is written in PHP, Python, Go, or Node.js, a "Zero Downtime" environment requires the application to follow the 12-factor principles.
Configuration and Environment Variables
In a dynamic cluster, you can't rely on .env files that change on disk. All configuration must be injected as Environment Variables from Kubernetes ConfigMaps and Secrets.
- Node.js/Python/Go: Read directly from the environment (
process.env/os.environ). - PHP: Ensure that configuration cached during build time doesn't contain hardcoded paths that change in production.
Queue Systems and Serialization
An often overlooked trap is asynchronous jobs (RabbitMQ, Redis, Kafka). When a new version is deployed, there may be jobs in the queue that are serialized with the old code structure. If a worker with new code picks up an old job payload, the application may crash.
Solution: Use versioned queues, or ensure that job payloads are always backward compatible. For major changes ("breaking changes"), the queue must be drained before updating.
12-Factor App Principles for ZDD
Never hardcode credentials or environment-specific values. Everything is injected via ConfigMaps and Secrets.
Each pod should be able to die and be replaced at any time. No local state on disk.
The application exposes itself via a port. No dependency on external web server.
Fast startup and graceful shutdown. Handle SIGTERM correctly.
Results and Business Value
Implementation of this architecture yielded the following results:
Deployment Frequency: Increased from monthly to 8.5 times per day. Developers now deploy small changes continuously.
Lead Time: Reduced from 5 days to 45 minutes (from commit to production).
Availability: Achieved 99.99% uptime the first year, even through major refactorings.
Culture: "Deployment fear" disappeared. Tuesday evening went from "on-call evening" to free time.
Conclusion
You don't get zero downtime "for free" just by choosing Kubernetes. It comes from a deliberate architecture: solid CI/CD pipelines, declarative infrastructure (GitOps), and a good understanding of the application's lifecycle.
For companies competing in a market that demands 24/7 availability, the investment in this architecture is a business advantage as much as an IT cost.
At PXL, we help companies modernize their deployment strategy. We have set up scalable, fault-tolerant environments for applications built in everything from PHP and Python to Go and Node.js, and we take the whole journey: containerization, CI/CD design, Kubernetes operations and monitoring.