EngineerJobs.io
← Back to all jobs

Job Description

CareView Communications, Inc. is hiring a Site Reliability Engineer to own reliability and operational health for on-premises servers and field-deployed systems in a hybrid environment.

Responsibilities

  • Design, implement, and maintain high availability architecture across on-premises and cloud-connected environments to support continuous uptime for clinical systems
  • Detect single points of failure across network, application, and data layers and create mitigation strategies
  • Implement zero-downtime deployment patterns, including blue/green and rolling updates, to remove planned maintenance windows
  • Partner with Development and Technical Operations teams to ensure high availability is addressed at every layer of the platform
  • Own the health, stability, and performance of a fleet of on-premises servers deployed at customer clinical sites
  • Establish and maintain remote monitoring and telemetry for real-time visibility into the behavior of field hardware
  • Create and run remote diagnostic and remediation procedures that minimize disruption to live clinical environments
  • Build and maintain a structured OS update and patch management program for fielded and cloud-based server infrastructure
  • Plan staged OS rollout strategies aligned with live clinical constraints, including customer change control windows and regulatory requirements
  • Maintain hardened OS baseline configurations in line with applicable security and regulatory frameworks
  • Design and operate a monitoring and alerting platform for system health, performance metrics, error rates, and availability
  • Develop automated health checks and self-healing to detect and recover from failures before end-user impact
  • Maintain operational runbooks to standardize incident response
  • Write code and automation that remove recurring manual operational tasks to reduce toil and improve scalability
  • Develop infrastructure-as-code for consistent, repeatable provisioning and configuration of cloud and on-premises resources
  • Design and maintain disaster recovery procedures, including regular DR testing with documented RTO and RPO targets
  • Lead capacity planning, forecasting infrastructure requirements ahead of product growth and customer site expansion
  • Perform other related duties as assigned

Requirements

  • 5+ years of experience in site reliability engineering or infrastructure engineering
  • Proven experience managing on-premises server infrastructure in addition to cloud environments
  • Hands-on experience designing and implementing high availability architecture, including load balancing, clustering, and automated failover in an on-premises environment
  • Strong Linux administration with command-line proficiency, including production hardening and maintenance
  • Proficiency with containerization and orchestration, including Docker and Kubernetes
  • Experience with monitoring and observability platforms such as Prometheus and Grafana
  • Bachelor’s degree in computer science or a related field, or equivalent practical experience

Technologies

  • Linux
  • Docker
  • Kubernetes
  • Prometheus
  • Grafana

Preferred Experience (Not Required)

  • Experience with VMware, Hyper-V, or other hypervisor platforms in production environments

Benefits

  • Health insurance
  • Dental insurance
  • Vision insurance
  • Life insurance
  • Paid time off
  • 7 paid holidays
  • Employee assistance program

Work Environment

  • Must pass a criminal background check and drug test
  • Must pass a HIPAA compliance test
  • Must be authorized to work in the United States
  • Working hours: 8:00 AM to 5:00 PM
  • Occasional evening and weekend work may be required as job duties demand
  • Hybrid
  • Ability to commute: Lewisville, TX 75067 (Required)
  • Work location: Hybrid remote in Lewisville, TX 75067
  • Office address: 405 State Hwy 121 Bypass, Suite B240 Lewisville, TX 75067

Pay

  • $95,000.00 - $110,000.00 per year

Similar Jobs