Senior SRE

TechGrove by Banyan Software · Bengaluru, India · Engineering

Posted 2026-07-13

Apply for this role →

Job Title: Senior SRE (Site Reliability Engineer) – Modernized Application Operations

Overview

We are seeking a highly experienced and hands-on SRE to own the operational excellence of the modernized SaaS applications produced by the Banyan AI Factory. This is not a role focused on building the factory itself; instead, you will run the reliability of the modernized applications the factory delivers to our Operating Companies (OpCos).

You will join a team that provides 24x7 coverage with rotating on-call responsibilities, serving as Tier 1 Site Reliability Engineering (SRE) for our OpCos’ distributed applications. Day to day this will include: automated deployments, cloud service integration, application performance and availability monitoring/observability, and security incident response across our two target clouds — Amazon Web Services (AWS) and Microsoft Azure. The ideal candidate has a track record of keeping secure, highly available production systems running at scale.

Key Responsibilities

24x7 Operations & On-Call: Operate as part of a team providing round-the-clock coverage of OpCo containerized applications, participating in a rotating on-call schedule to ensure continuous availability and rapid response.

Tier 1 SRE & Operations: Serve as Tier 1 SRE for the modernized applications, managing day-to-day cloud integrations across our two target clouds — AWS and Azure — to keep production systems healthy, performant, and secure.

Performance & Availability Monitoring/Observability: Implement and maintain robust application observability tooling (monitoring, logging, tracing) to track performance and availability, proactively detect degradation, and drive down mean-time-to-detect and mean-time-to-resolve.

Security Incident Response: Respond to security incidents and operational events affecting OpCo SaaS platforms, executing established runbooks, coordinating remediation

Automation & Infrastructure-as-Code : Use Infrastructure-as-Code (Terraform) and CI/CD pipelines (e.g., GitHub Actions, GitLab CI) to manage, deploy, and automate the operational environments of modernized applications, reducing toil and improving consistency.

AI Agents & DevSecOps Scale: Build scale in our DevSecOps practice by designing, building, and operating AI agents that automate SRE tasks and incident response, reducing toil and accelerating detection, triage, and remediation.

Hands-on Problem Solving: Serve as a technical escalation point for operational challenges, applying strong analytical skills to resolve infrastructure, network, and automation issues across distributed, multi-tenant SaaS environments while navigating technical ambiguity.

Required Qualifications & Experience

Experience: 5–7 years of progressive experience in Software Engineering, and/or Site Reliability Engineering, with a focus on operating distributed systems.

Automation Coding Experience: Deep expertise in Python, Javascript, or Go. Building automation and integrations between tools. This may be with AI assistance, but you must have a deep understanding of the code and scripting principals such as: authentication, parallelization, triggering, APIs, data transformation, etc.

Containerization: Deep expertise in container technologies (Docker/Kubernetes) supporting highly scalable and resilient distributed systems.

Infrastructure-as-Code with Terraform: Have experience working with modules at scale. This is a requirement for the role.

Cloud Native Services: hands-on experience operating production workloads on Amazon Web Services (AWS) (e.g., EC2, Lambda, EKS, S3, RDS) and / or Microsoft Azure (e.g., Container Apps, AKS, Container Storage).

CI/CD & Automation: Deep history of hands-on work with CI/CD platforms (GitHub Actions, GitLab CI) and embedding DevSecOps practices directly into operational workflows.

Operations, Monitoring & Observability: Experience with application level logging, troubleshooting, and tracing tools, with a proven track record operating highly available production systems.

AI-Fluent Engineering: Experience with AI-assisted engineering tools such as Claude Code or similar

Application Performance Management (APM): Familiarity with APM tooling and practices (e.g., Datadog, New Relic, Dynatrace, or similar) to instrument, profile, and optimize application performance in production.

Incident & Security Response: Demonstrated experience participating in on-call rotations, responding to production and security incidents, and executing disaster recovery procedures.

Communication & Collaboration: Exceptional communication, presentation, and collaboration skills, with a proven ability to coordinate across teams.

Education: Bachelor’s degree in Computer Science or a related technical field.

Preferred Skills (A Plus)

Familiarity with advanced cloud security tools like Wiz, Prisma Cloud, and Checkov.

Apply for this role →

← Back to all jobs