Microsoft Hiring Site Reliability Engineer 2 (SRE) – Azure, Python, PowerShell, AI Automation, Incident Management & Cloud Operations – Hyderabad, India

Job Description
Microsoft is hiring a Site Reliability Engineer 2 (SRE) to join its Azure Data Engineering organization in Hyderabad. This role is part of the Customer Data Integration (CDI) team within Microsoft Fabric and focuses on ensuring the reliability, availability, and operational excellence of Microsoft's large-scale cloud services.
As an SRE, you will be responsible for incident triage, production operations, live site management, automation, observability, and cloud service reliability. You will provide first-line support for critical services including Microsoft Fabric, Power BI, Power Query Online, Azure Data Factory, Azure SQL Database, Azure Cosmos DB, Azure PostgreSQL, Azure Service Bus, Azure Event Grid, and Azure Synapse Analytics.
The role combines traditional Site Reliability Engineering with AI-powered automation. You will develop intelligent agents that analyze incidents, correlate alerts with deployments and feature flag rollouts, identify known issues, recommend incident severity, automate routing, generate customer communications, create postmortems, and improve operational efficiency.
You will work closely with software engineering, cloud infrastructure, and support teams to reduce operational toil, improve service availability, automate repetitive tasks, and enhance incident response processes. This position includes participation in an on-call rotation supporting business-critical cloud services across global regions.
This is a full-time hybrid role, requiring employees to work three days per week from the Hyderabad office.
Requirements
- Minimum 4 years of experience in Site Reliability Engineering (SRE), Live Site Operations, Incident Management, or Cloud Software Engineering.
- Strong programming skills in Python, C#, PowerShell, or Kusto Query Language (KQL).
- Experience with incident management platforms such as ICM, PagerDuty, ServiceNow, or similar tools.
- Experience with monitoring, alerting, and observability platforms including Grafana, Kusto, Geneva, or equivalent.
- Strong understanding of cloud operations, production support, and service reliability.
- Experience participating in 24×7 on-call rotations and handling production incidents.
- Knowledge of incident lifecycle management, root cause analysis, and operational excellence.
- Ability to build automation for incident response, alert correlation, and operational workflows.
- Strong troubleshooting, debugging, and analytical skills.
- Experience collaborating with engineering teams, leadership, customer support, and stakeholders.
- Understanding of SLA management, customer communications, and cloud service escalation processes.
- Experience with Azure Cloud, Microsoft Fabric, Power BI, or related Microsoft cloud services is preferred.
- Knowledge of AI-driven automation, LLMs, Copilot extensibility, MCP servers, or agentic AI frameworks is preferred.
- Experience with telemetry analysis, log analysis, troubleshooting guides (TSGs), and incident pattern analysis is an advantage.
- Ability to meet Microsoft's security screening and Cloud Background Check requirements.