Site Reliability Engineer
Analytics
Cloud
Cloud Infrastructure
Cloud Native
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Technology
Dashboards
Data Analytics
Data Platform
Design
Desktop Support
DevOps
Devops Tools
DevSecOps
Docker
Embedded System
End User Support
Engineering
Grafana
High Availability
Information Technology (IT)
Infrastructure
Infrastructure As Code
Kubernetes Ops
Linux
Monitoring
Observability
Operating System
Operating Systems
Platform Engineering
Prometheus
Release Engineering
Reporting and Analytics
Security Automation
Site Reliability Engineering
Software Development
Systems
Technical Support
Telemetry
Unix Operating System
Visual Design
Job Description
CareView Communications, Inc. is hiring a Site Reliability Engineer to own reliability and operational health for on-premises servers and field-deployed systems in a hybrid environment.
Responsibilities
- Design, implement, and maintain high availability architecture across on-premises and cloud-connected environments to support continuous uptime for clinical systems
- Detect single points of failure across network, application, and data layers and create mitigation strategies
- Implement zero-downtime deployment patterns, including blue/green and rolling updates, to remove planned maintenance windows
- Partner with Development and Technical Operations teams to ensure high availability is addressed at every layer of the platform
- Own the health, stability, and performance of a fleet of on-premises servers deployed at customer clinical sites
- Establish and maintain remote monitoring and telemetry for real-time visibility into the behavior of field hardware
- Create and run remote diagnostic and remediation procedures that minimize disruption to live clinical environments
- Build and maintain a structured OS update and patch management program for fielded and cloud-based server infrastructure
- Plan staged OS rollout strategies aligned with live clinical constraints, including customer change control windows and regulatory requirements
- Maintain hardened OS baseline configurations in line with applicable security and regulatory frameworks
- Design and operate a monitoring and alerting platform for system health, performance metrics, error rates, and availability
- Develop automated health checks and self-healing to detect and recover from failures before end-user impact
- Maintain operational runbooks to standardize incident response
- Write code and automation that remove recurring manual operational tasks to reduce toil and improve scalability
- Develop infrastructure-as-code for consistent, repeatable provisioning and configuration of cloud and on-premises resources
- Design and maintain disaster recovery procedures, including regular DR testing with documented RTO and RPO targets
- Lead capacity planning, forecasting infrastructure requirements ahead of product growth and customer site expansion
- Perform other related duties as assigned
Requirements
- 5+ years of experience in site reliability engineering or infrastructure engineering
- Proven experience managing on-premises server infrastructure in addition to cloud environments
- Hands-on experience designing and implementing high availability architecture, including load balancing, clustering, and automated failover in an on-premises environment
- Strong Linux administration with command-line proficiency, including production hardening and maintenance
- Proficiency with containerization and orchestration, including Docker and Kubernetes
- Experience with monitoring and observability platforms such as Prometheus and Grafana
- Bachelor’s degree in computer science or a related field, or equivalent practical experience
Technologies
- Linux
- Docker
- Kubernetes
- Prometheus
- Grafana
Preferred Experience (Not Required)
- Experience with VMware, Hyper-V, or other hypervisor platforms in production environments
Benefits
- Health insurance
- Dental insurance
- Vision insurance
- Life insurance
- Paid time off
- 7 paid holidays
- Employee assistance program
Work Environment
- Must pass a criminal background check and drug test
- Must pass a HIPAA compliance test
- Must be authorized to work in the United States
- Working hours: 8:00 AM to 5:00 PM
- Occasional evening and weekend work may be required as job duties demand
- Hybrid
- Ability to commute: Lewisville, TX 75067 (Required)
- Work location: Hybrid remote in Lewisville, TX 75067
- Office address: 405 State Hwy 121 Bypass, Suite B240 Lewisville, TX 75067
Pay
- $95,000.00 - $110,000.00 per year