Resources Blog
Blog

Platform engineering, in the open.

Field notes on running Kubernetes with less toil — AI SRE and incident response, CI/CD and golden paths, continuous security, cost, and the internal developer platform that ties them together.

AI SRE for Kubernetes: The Complete GuideFeatured
AI SRE

AI SRE for Kubernetes: The Complete Guide

AI SRE means an agent that carries a Kubernetes incident from detection through diagnosis to a proposed fix, then checks whether the fix worked. This guide covers the full loop, the permission model, what to evaluate in any tool, and where the category still falls short.

Aug 2026AI SRE
AI SRE Security: Safe AI Automation for Production KubernetesAI SRE

AI SRE Security: Safe AI Automation for Production Kubernetes

Every AI SRE evaluation reaches the same objection: you want an AI to have production access? This is the honest answer — what to grant, what to refuse, the RBAC to demand instead of cluster-admin, and the audit record that gets it past a security review.

AI SRE for Amazon EKS: Automating Kubernetes Operations on AWSAI SRE

AI SRE for Amazon EKS: Automating Kubernetes Operations on AWS

The EKS failures that cost the most time are not Kubernetes failures. VPC CNI IP exhaustion, IRSA trust policy mismatches and add-on version drift all present as ordinary pod problems while their causes sit outside the cluster. Here is what an agent handles and what it does not.

How to Reduce MTTR in Kubernetes: 6 Proven StrategiesAI SRE

How to Reduce MTTR in Kubernetes: 6 Proven Strategies

On Kubernetes, MTTR is dominated by diagnosis, not the fix. Six proven strategies to reduce Kubernetes MTTR — unified telemetry, incident grouping, AI root-cause analysis, and automated GitOps remediation.

Best Internal Developer Platforms in 2026: 9 ComparedPlatform Engineering

Best Internal Developer Platforms in 2026: 9 Compared

Nine internal developer platforms compared honestly — Backstage, Port, Cortex, OpsLevel, Humanitec, Qovery, Northflank, Cycloid and Atmosly — organised by the one distinction that matters: portals that organise vs execution layers that ship.

GitOps Remediation for Kubernetes: From Incident to Merged PRAI SRE

GitOps Remediation for Kubernetes: From Incident to Merged PR

A live kubectl patch on a GitOps cluster gets reverted at the next sync. GitOps remediation fixes incidents in Git — the automated alert-to-PR pattern, and how to keep it safe.

Kubernetes Node NotReady: Causes, Diagnosis & FixesKubernetes

Kubernetes Node NotReady: Causes, Diagnosis & Fixes

A NotReady node stops scheduling and can evict your pods. This guide covers what NotReady means, the exact kubectl commands to diagnose it, and the fix for every cause — kubelet, runtime, resource pressure, CNI, and certificates.

Kubernetes Audit Logging: Policy, Setup, and Retention for Compliance (2026)Security

Kubernetes Audit Logging: Policy, Setup, and Retention for Compliance (2026)

Kubernetes audit logging end to end: the four audit levels, a production audit policy YAML, enabling logs on EKS (off by default), GKE, and AKS, retention that satisfies SOC 2, PCI DSS, and HIPAA, and the five alerts worth wiring up.

Kubernetes Security & Compliance Platforms Compared (2026)Security

Kubernetes Security & Compliance Platforms Compared (2026)

Wiz, Prisma Cloud, Aqua, Sysdig, ARMO/Kubescape, and Atmosly compared honestly — by the three jobs Kubernetes security platforms actually do (posture, CVEs, runtime), with a coverage table and a decision path to your shortlist.

Kubernetes Compliance: Frameworks, Controls, and Evidence (2026)Security

Kubernetes Compliance: Frameworks, Controls, and Evidence (2026)

Kubernetes compliance, end to end: what CIS, SOC 2, ISO 27001, PCI DSS, and HIPAA actually require of a cluster, the shared responsibility line on managed Kubernetes, the six control families that satisfy every framework at once, and the scan-fix-evidence loop that passes audits.

Mapping Kubernetes Controls to ISO 27001 Annex A: A Practical GuideSecurity

Mapping Kubernetes Controls to ISO 27001 Annex A: A Practical Guide

ISO 27001 never mentions Kubernetes, so the clause-to-control mapping is on you. This guide covers the ten Annex A clauses a cluster scan can evidence, the four it cannot, and a full walkthrough of A.8.22.

CIS Kubernetes Benchmark: A Practical Implementation Guide (2026)Security

CIS Kubernetes Benchmark: A Practical Implementation Guide (2026)

First CIS scans land in the 60s because Kubernetes defaults favour startup over containment. Here is how to run the benchmark, remediate the four checks every fresh EKS cluster fails, and verify each fix.

Less reading about toil. Less toil.

Connect a cluster read-only and see your incidents, spend, and deploys in one place — in minutes. Free, no sales call.

Connect your cluster → See the platform