DevOps / Platform Infrastructure Engineer
Description
COPE Health Solutions is hiring a Senior DevOps / Platform Infrastructure Engineer to help run and improve the platform behind our applications. Our infrastructure lives primarily in Microsoft Azure. We deploy through GitHub Actions, run workloads on Kubernetes, manage multiple databases, and support systems that handle healthcare data. This is a hands-on role. You'll spend more time in Bicep templates, AKS clusters, GitHub Actions workflows, PowerShell scripts, and architecture diagrams than you will in meetings about them. You'll also be one of the people developers call when a deployment breaks at 4:30 on a Friday afternoon. Patience, curiosity, and a healthy sense of humor go a long way. Because we work with healthcare data, security, compliance, and operational discipline aren't side projects. They're part of the job.
FLSA Status | Exempt | Salary Range | $128,600-$181,600 |
Reports To | Director of Data Information and Security | Direct Reports | Yes |
Location | Remote | Travel | Up to 10% |
Work Type | Regular | Schedule | Full Time |
Position Description:
- Build and maintain Azure infrastructure using Bicep
- Operate and improve our AKS environments, including upgrades, scaling, reliability, and cost optimization
- Own CI/CD pipelines built with GitHub Actions
- Automate repetitive operational work using PowerShell and Python
- Implement monitoring, logging, and alerting that helps us discover problems before users do
- Reduce cloud spend without creating new operational risks
- Support PostgreSQL, SQL Server, and MySQL environments, including backups, maintenance, performance tuning, and disaster recovery planning
- Design and document network architecture, including hub-and-spoke VNets, private endpoints, subnet segmentation, peering, routing, and hybrid connectivity
- Support identity, access management, secrets management, and compliance initiatives
- Improve developer experience through better tooling, documentation, automation, and sensible platform defaults
- Participate in incident response, root-cause analysis, and operational reviews
Documentation is part of the job
We expect documentation to be treated like code.
You'll be responsible for creating and maintaining:
- Runbooks
- Architecture diagrams
- Onboarding guides
- Incident postmortems
- Operational procedures
- Platform documentation
If someone asks how traffic gets from Point A to Point B, there should be a diagram that answers the question.
Qualifications:
What we're looking for
- 5+ years operating production infrastructure
- Deep hands-on experience with Microsoft Azure
- Strong Infrastructure-as-Code experience with Bicep
- Extensive Networking experience
- Production Kubernetes experience, preferably AKS
- Strong GitHub Actions and CI/CD experience
- Strong PowerShell and Python scripting skills
- Experience supporting Linux and Windows Server environments
- Experience supporting PostgreSQL, and SQL Server
- Experience designing and troubleshooting network architectures
- Experience with observability platforms such as Azure Monitor, or similar tools
- Experience implementing RBAC, IAM, secrets management, and security controls
- Experience working in regulated environments such as HIPAA,HiTrust, and SOC 2
- Strong written communication and documentation skills
On-call and production support
We run systems clients and their members depend on, so someone needs to be reachable when something breaks outside business hours. We try to be humane about how we do that.
Platform engineers share a primary/secondary on-call rotation. Daytime issues get picked up by whoever's around; the rotation exists for evenings, weekends, and holidays. It covers genuine production incidents: a service down, a data-path failure, a security event, not routine tickets or "can you look at this sometime" requests. When you're paged, you're paged for something that actually matters.
We expect platform engineers to participate in incident response, troubleshooting, and root-cause analysis. Just as importantly, we expect recurring operational pain to be addressed through automation, monitoring, documentation, or engineering improvements. The goal isn't to become better at responding to the same alert every week; it's to make sure that alert stops happening.
We've all been on the receiving end of bad on-call rotations. We'd rather invest in reliability than heroics.
What success looks like
Month 1
- Get access
- Learn the environment
- Ship something small
- Join incident reviews
- Figure out where the sharp edges are
Month 2
- Take ownership of one major platform area
- Contribute meaningful automation or operational improvements
- Start participating in production support activities
Month 3
- Identify risks we haven't seen yet
- Propose improvements we haven't thought of
- Help raise the engineering standard of the platform
Interview process
We're less interested in trivia than in how you think.
Candidates should expect practical technical discussions, including:
- Reviewing a Bicep template
- Debugging a GitHub Actions pipeline
- Diagnosing an AKS issue
- Explaining a hub-and-spoke Azure network design
- Discussing a production incident they've personally handled
- Walking through a security or HIPAA-related scenario
A strong answer doesn't require perfection. We're looking for engineers who can reason through problems, communicate clearly, and learn from operational experience.
How we work
You'll join a team of 13 and work closely with engineering and security teams. Code review is required. Ego is optional. Documentation gets reviewed like code. When incidents happen, we focus on fixing root causes rather than assigning blame.
Benefits:
As a firm passionate about health care, we’re deeply committed to the health and wellness of our own team members. We offer comprehensive, affordable insurance plans for our team and their families, and a host of other unique benefits, such as a yearly stipend for wellness-related activities, and a paid parental leave program. You can learn more about our benefits offerings here: https://copehealthsolutions.com/careers/
About COPE Health Solutions
COPE Health Solutions is a national tech-enabled services firm powering success for health plans and for providers in risk arrangements. Our comprehensive NCQA certified population health management platform and highly experienced team brings deep expertise, experience, proven tools, and processes to improve financial performance and quality outcomes for all types of payers and providers. CHS de-risks the roadmap to advanced value-based payment and improves quality and financial performance for providers, health plans and self-insured employers. For more information, visit https://copehealthsolutions.com/about-us/
To Apply:
To apply for this position or for more information about COPE Health Solutions, visit us at https://copehealthsolutions.com/careers/open-positions/