Friday, September 18, 2026
CLOUDSUTRA

Vegapay is Hiring Site Reliability Engineer (SRE)

VegapayBengaluru, Karnataka, India
INR ₹8,00,000 - ₹15,00,000/year
Vegapay
DevOpsFull-timePosted 7 June 2026Updated 7 June 2026Expired
Share:

Job Description

Vegapay is building modern fintech infrastructure that enables banks, NBFCs, and enterprises to design, launch, and manage credit and payment programs with speed and flexibility. Our platform helps financial institutions overcome legacy system limitations and deliver scalable, personalized financial experiences.

We are focused on creating configurable and high-performance infrastructure that powers next-generation fintech innovation.

About the Role

We are looking for a proactive and detail-oriented Site Reliability Engineer (SRE) to join our production operations and reliability engineering team. In this role, you will be responsible for ensuring the reliability, availability, and performance of mission-critical systems and infrastructure.

You will work closely with engineering and operations teams to monitor production environments, respond to incidents, oversee scheduled processing jobs, and improve operational efficiency through automation and observability best practices.

Key Responsibilities

Production Monitoring & Reliability

  • Monitor application health checks and maintain high system availability
  • Monitor infrastructure health including servers, databases, cloud resources, and networking components
  • Continuously track Grafana dashboards to identify anomalies, traffic spikes, and performance bottlenecks
  • Review and respond to alerts generated from monitoring and observability platforms

Incident Management & Support

  • Create, manage, and track production incidents through resolution
  • Perform initial troubleshooting and coordinate with engineering teams for issue resolution
  • Follow incident escalation procedures and SLA guidelines
  • Maintain incident logs, RCA documentation, and operational runbooks
  • Participate in on-call rotations and support critical production activities

Batch Processing & Scheduled Jobs

  • Monitor end-of-day batch processing workflows and scheduled cron jobs
  • Ensure successful execution of reporting jobs and operational processes
  • Troubleshoot failed scheduled tasks and coordinate recovery actions

Observability & Automation

  • Work with monitoring tools such as Grafana, Prometheus, CloudWatch, and related platforms
  • Improve observability practices across infrastructure and applications
  • Identify opportunities to automate repetitive operational tasks
  • Contribute to reliability improvements and operational excellence initiatives

Why Join Vegapay?

  • Opportunity to work in a fast-growing fintech environment
  • Exposure to large-scale production systems and cloud infrastructure
  • Collaborative engineering culture focused on innovation and reliability
  • Hands-on experience with observability, automation, and cloud-native operations
  • Career growth opportunities in SRE, DevOps, and Platform Engineering
  • Work on impactful financial technology products used at scale

Requirements

  • 2+ years of experience in Site Reliability Engineering, Production Support, DevOps, or Infrastructure Operations
  • Hands-on experience with monitoring and observability tools such as Grafana, Prometheus, CloudWatch, or similar platforms
  • Familiarity with AWS, Azure, or Google Cloud Platform environments
  • Experience monitoring scheduled jobs, cron-based workflows, and batch processing systems
  • Strong understanding of incident management and production support processes
  • Ability to troubleshoot infrastructure, application, and system performance issues
  • Knowledge of Linux/Unix systems and operational best practices
  • Understanding of networking, servers, databases, and cloud infrastructure concepts
  • Experience maintaining operational documentation and runbooks
  • Ability to work effectively in a fast-paced, high-availability production environment
  • Good analytical, debugging, and problem-solving skills
  • Strong written and verbal communication abilities
  • Experience with automation or scripting is a plus
  • Familiarity with DevOps and cloud-native operational practices is an advantage