Years of Experience 5–8 Years
Design, implement, and manage highly available, scalable, and secure cloud
infrastructure on AWS.
Build and maintain an end-to-end observability platform using Open
Telemetry, Grafana, Datadog, CloudWatch, and related tools.
Implement AIOps capabilities, including:
LLM-assisted incident triage
AI-powered root cause analysis
ML-driven forecasting and anomaly detection
Intelligent alert correlation and noise reduction
Lead production incident management, on-call response, postmortems, and
Root Cause Analysis (RCA).
Automate operational workflows using Infrastructure as Code
(Terraform/CloudFormation) and CI/CD pipelines.
Drive infrastructure rightsizing, capacity planning, utilization analysis, and
cloud cost optimization.
Build dashboards, SLOs, SLIs, and error budgets to improve service reliability.
Develop automation scripts using Python, Bash, or Go to eliminate manual
operational tasks.
Monitor application and infrastructure health while ensuring high uptime and
service performance.