About the author
About
Site reliability engineer writing practical DevOps and SRE material from production experience.
Who I am
I write as SBM SRE. I am a site reliability engineer with more than thirteen years in infrastructure and operations. I work on production platforms running many services across cloud and hybrid environments, with 24/7 operations and strict compliance requirements.
I built this site to share what I use and teach: practical, tested procedures, and learning paths that go from the basics to on-call readiness.
How I approach production
- Fail fast on systemic problems instead of retrying blindly.
- Prefer automatic, resumable recovery over manual intervention.
- Avoid collateral damage: a fix must not create a problem for a dependent system.
- Change one thing at a time, and record it.
Areas of focus
- Site reliability engineering and incident management
- Kubernetes and cloud platforms
- Infrastructure as Code and automation
- Observability and alerting
Tools I work with
AWSKubernetes / OKDTerraformAnsiblePackerArgo CDGitLab CI/CDDockerDatadogPrometheusELKPythonBashNexus
Where to start
benmabrouk.fr: free DevOps and SRE learning resources, written from production experience.