Cloud Site Reliability Senior Engineer
Come Join Our Passionate Team! At Barracuda, we make the world a safer place. We believe every business deserves access to cloud-enabled, enterprise-grade security solutions that are easy to buy, deploy, and use. We protect email, networks, data, and applications with innovative solutions that grow and adapt with our customers’ journey. More than 200,000 organizations worldwide trust Barracuda to protect them — in ways they may not even know they are at risk — so they can focus on taking their business to the next level.
We know a diverse workforce adds to our collective value and strength as an organization. Barracuda Networks is proud to be an Equal Opportunity Employer, committed to equal employment opportunity and equitable compensation regardless of race, gender, religion, sex, sexual orientation, national origin, or disability.
Envision yourself at Barracuda
Barracuda Networks is looking for a Cloud Site Reliability Senior Engineer to join the centralized Cloud Operations team. In this role, you will help lead reliability improvements for production services, partner closely with Engineering and Operations teams, and contribute to scalable, resilient, and secure service delivery across Barracuda. The role includes a focus on Data Inspector while also encouraging shared ownership, cross-product collaboration, and knowledge sharing across SRE teams.
What you'll be working on
- Help improve the reliability, availability, and operational excellence of production services through shared team ownership.
- Lead and contribute to troubleshooting, incident response, RCA, and preventive improvements across cloud services and supporting platforms.
- Partner with Engineering, Cloud Operations, and SRE teams to improve release readiness, deployment practices, and operational standards.
- Support cloud platform improvements, automation, and resilience initiatives that reduce toil and improve service health.
- Strengthen observability, alerting, and operational insights so teams can respond faster and improve service reliability.
- Share knowledge, improve documentation, and help reduce product or domain silos across SRE and Cloud Operations teams.
- Contribute to secure, compliant, and cost-conscious cloud operations across Azure, AWS, Terraform-managed infrastructure, and containerized services.
What you bring to the role
- Bachelor’s degree in computer science engineering, Information Technology, or equivalent degree.
- 5+ years of progressive experience in Site Reliability Engineering, Cloud Operations, DevOps, or Platform Engineering.
- Strong Linux/Unix command-line administration, troubleshooting, package management, and systems operations skills.
- Experience managing production cloud infrastructure across Azure and/or AWS environments.
- Hands-on experience operating container orchestration platforms, including AKS migrations, upgrades, and production operations.
- Hands-on Infrastructure as Code experience, preferably with Terraform for provisioning and lifecycle management.
- Automation, scripting, and AI-assisted engineering experience using Python, Bash, Go, YAML, generative AI tools, or similar technologies.
- Experience with CI/CD pipelines using Azure DevOps, ArgoCD, Jenkins, GitHub Actions, or similar platforms.
- Experience with Docker, container registries such as Azure Container Registry, and container lifecycle management.
- Experience with configuration management and automation tools such as Ansible, Puppet, or Chef.
- Experience operating monitoring, observability, and incident response platforms such as PagerDuty, Grafana, Prometheus, ELK/OpenSearch, New Relic, or Sensu.
- Track record of practical problem solving, RCA, and corrective actions that improve service reliability.
- Strong understanding of networking, the OSI model, SQL/NoSQL databases, and troubleshooting distributed systems.
- Ability to prioritize tasks, work independently, and communicate clearly with technical and non-technical audiences.
- Enthusiasm for collaborating with globally distributed teams through video conferencing, Slack, and other communication tools.
Preferred Qualifications
- Experience contributing to large-scale AKS migrations, container platform modernization, or cloud transformation programs.
- Proven ability to coordinate complex production releases across Engineering, Operations, and stakeholder teams.
- Exposure to AI-assisted operations, generative AI tools, or automation use cases that improve engineering productivity.
- Experience building reliability tooling, self-healing capabilities, or automation frameworks that reduce operational toil.
- Relevant certifications in Azure, AWS, Terraform, container platforms, or cloud reliability engineering.
What you’ll get from us
A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility – there are many opportunities for cross training and the ability to define your next career step within Barracuda. We support employees who want to explore other areas of interest.