About – The Production Engineering Digest
The Production Engineering Digest is a practical newsletter for engineers who build, support, and operate live production systems.
Production is where theory meets reality. Systems behave unexpectedly, incidents happen at inconvenient hours, and decisions must be made under pressure. This newsletter focuses on what actually works when software is running in the real world.
Each edition breaks down:
Real production incidents and postmortems
Reliability and stability patterns
Production support and operations best practices
Lessons learned from on‑call, war rooms, and live outage recovery
Career insights for engineers working close to production
This is not about perfect architectures or ideal workflows. It’s about handling failure gracefully, maintaining uptime, and continuously improving systems that users depend on every day.
Why subscribe?
Subscribe to get full access to the newsletter and publication archives.
Who This Is For?
This newsletter is written for:
Production Support Engineers (L1 / L2 / L3)
Application Support & Operations Engineers
Site Reliability Engineers (SREs)
DevOps Engineers
Technical Leads and Engineering Managers
Engineers transitioning from development to production roles
If you’ve ever debugged an issue with incomplete logs, joined a late‑night bridge call, or been responsible for keeping systems running—you’re in the right place.
What You’ll Get
Subscribers can expect:
Clear explanations of complex production issues
Practical approaches to incident response and root cause analysis
Realistic guidance on monitoring, alerting, and reliability
Honest insights into production engineering careers
Content grounded in enterprise, high‑availability systems
The goal is simple: help engineers think better under pressure and build systems that stay online.
About the Author
Arun Sankar A S K is a Software Engineer specializing in Production Engineering and Application Production Support, with over three years of experience supporting enterprise‑grade systems in 24/7 live production environments. He works as an L2 Production Support and Maintenance Engineer, supporting post‑trade derivatives settlement and confirmation applications for a global financial services client.
His experience spans end‑to‑end incident management, detailed root cause analysis (RCA), problem management, service request fulfillment, and change and release support within ITIL‑driven environments. Arun has handled business‑critical production incidents, performed impact analysis, and ensured timely service restoration while meeting strict SLA and reliability targets.
He works extensively with SQL, PL/SQL, UNIX, and Shell scripting, production monitoring using Geneos and Splunk, and batch scheduling via Control‑M. He contributes to automation and operational improvements, focusing on reducing manual intervention and improving system stability in production environments.
Arun is deeply involved in knowledge management and operational readiness, creating and maintaining SOPs, runbooks, troubleshooting documentation, and knowledge articles that enable faster incident resolution and smoother onboarding for support teams. He also participates in production deployments, disaster recovery activities, and live environment validations, ensuring systems remain stable during critical changes.
Through The Production Engineering Digest, Arun shares practical lessons from real production systems—covering incidents, reliability, operational discipline, and day‑to‑day production engineering practices for engineers working close to live systems.
Join the crew
Be part of a community of people who share your interests. Participate in the comments section or support this work with a subscription.
To learn more about the tech platform that powers this publication, visit Substack.com.


