Site Reliability Engineer (SRE) Job at Halo Media, United States

  • Halo Media
  • United States

Job Description

About the Role:

We are looking for a Senior Site Reliability Engineer (SRE) to help modernize large-scale infrastructure and improve the reliability, scalability, and operational excellence of critical production systems. In this role, you will lead OS modernization initiatives, drive infrastructure automation, strengthen observability, and partner closely with engineering teams to support cloud migration efforts.

The ideal candidate has a strong background in Linux systems, Python, automation, CI/CD, and production operations, with a passion for building resilient platforms and solving complex infrastructure challenges.

Responsibilities:

  • Lead large-scale operating system modernization projects, including migrations from RHEL7 to EL8/9 across approximately 1,700 systems and virtual machines .

  • Drive infrastructure and packaging migrations, including Chef to CINC and yinst to RPM .

  • Build, maintain, and configure RPM packages to support modern infrastructure deployments.

  • Develop automated operational runbooks and infrastructure automation to improve efficiency and reliability.

  • Harden CI/CD pipelines, rollout/rollback mechanisms, and deployment processes for infrastructure modernization.

  • Strengthen observability by onboarding services to modern monitoring and logging platforms.

  • Triage, investigate, and resolve complex production incidents and critical software bugs.

  • Provide Tier-2 operational support in a follow-the-sun model alongside Site Reliability Engineering and Cloud Infrastructure teams.

  • Support incident response, troubleshooting, and break/fix activities across distributed production environments.

  • Partner with software engineering teams during cloud migration initiatives, providing operational guidance and technical support.

  • Automate repetitive operational tasks and maintain comprehensive technical documentation.

  • Drive reliability, operational excellence, and continuous improvement across production systems.

Required Qualifications:

  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Software Engineering with a strong infrastructure focus.

  • Strong hands-on experience with Python .

  • Proven experience leading Linux operating system modernization and infrastructure migration projects.

  • Experience with RHEL , Linux administration, and package management.

  • Hands-on experience building and maintaining RPM packages .

  • Strong experience with infrastructure automation and configuration management.

  • Experience designing, maintaining, and improving CI/CD pipelines.

  • Strong troubleshooting and incident management skills in large-scale production environments.

  • Experience supporting distributed systems with a strong focus on reliability and availability.

  • Excellent scripting, automation, and problem-solving skills.

Benefits:

  • 100% Remote.

  • International and collaborative environment.

Job Tags

Remote work

Similar Jobs

All Star Healthcare Solutions

Physician (MD/DO) - Cardiology - Interventional in South Carolina Job at All Star Healthcare Solutions

 ...Doctor of Medicine | Cardiology - Interventional Location: South Carolina Employer: All Star Healthcare Solutions Pay: Competitive weekly pay (inquire for details) Start Date: ASAP About the Position Experienced Interventional Cardiologist position... 

Bolivar-Richburg Central School District

Maintenance Mechanic Job at Bolivar-Richburg Central School District

 ...Maintenance Position Under the direct supervision of the Director of Facilities, to maintain school buildings and building and grounds in orderly, neat, safe and operable condition. Essential job functions include performing tasks to keep school buildings and building... 

Neshent Technologies

Site Reliability Engineer Job at Neshent Technologies

 ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable, scalable, secure, and highly available...  ...and data platforms. The role combines database engineering, site reliability engineering, Linux systems administration, and infrastructure... 

Virtual Vocations Inc

Institutional Giving Manager Job at Virtual Vocations Inc

 ...Experienced grant writers will find a full-time remote opportunity as an Institutional Giving Manager, responsible for managing the...  ...experience in grant writing and managing institutional grants in a nonprofit contextProven track record of managing the full grant... 

Conservation Corps North Carolina

Climate Forestry Job at Conservation Corps North Carolina

 ...Chainsaw Crew Leader - Climate Impact MitigationCrew Reports to: Program Manager Locations: Based out of Bahama, NC but will be camping and completing service projects on either North Carolinas Outer Banks or Croatan National Forest Season Dates: January...