HPC Data Center Production Engineer (Automation) - Banking & Finance
- $175,000 – $235,000 + Performance-Based Bonus (Commensurate with experience)
- Chicago, Illinois, United States
- Permanent
- 200000
- Artificial Intelligence
- AI Data Center
Are you looking for an exciting new opportunity?
Join a specialist provider of sophisticated point-to-point wireless networks supporting mission-critical trading operations across multiple financial markets, with a strong focus on quality, safety, and technical excellence.
The organization is currently on the lookout for an HPC Data Center Production Engineer to build, own, and scale the automation and software systems powering its high-performance computing data centres. The ideal candidate will transform raw hardware servers, switches, rack PDUs, CDUs, and liquid cooling systems into highly automated, production-ready infrastructure, translating operational strategies into software, designing outage simulations, and building advanced telemetry pipelines. Strong Linux fluency and clean Go coding ability are essential for this development-heavy, hands-on role.
Ready to make a move? Get in touch and apply today!
Responsibilities:
- Hardware Onboarding Automation: Design and build end-to-end automated workflows that take servers, switches, PDUs, CDUs, and environmental sensors from racked-and-cabled to production-ready.
- Capacity & Simulation Tooling: Develop tools for power/cooling capacity planning, predictive utilization modeling, and outage simulations to test facility redundancy.
- Observability & Telemetry Integration: Build custom telemetry integrations and metrics pipelines (IPMI/Redfish, SNMP) to normalize data from colocation facilities into centralized observability platforms.
- Cross-Functional Architecture: Work directly with HPC Planning, Engineering, and Operations leads to convert manual pain points into maintainable, automated systems.
- Reliability & Maintenance: Own the full software lifecycle of all internal tools, participating in scheduled maintenance windows and performing root-cause analysis on failures.
- AI Tooling Acceleration: Leverage AI tools daily for code generation, data analysis, debugging, and predictive capacity modeling.
Skills/Must Have:
- Experience: 5+ years in Production Engineering, Site Reliability Engineering (SRE), or Infrastructure Automation within HPC or large-scale data center environments.
- Programming Proficiency: High proficiency in Golang alongside strong scripting capabilities in Python.
- Linux Expertise: Mastery of Linux systems administration, OS-level troubleshooting, networking, and process management.
- Hardware & Data Center Domain: Deep understanding of data center power/cooling infrastructure (air and liquid), structured cabling, and hardware management interfaces (IPMI, BMC, Redfish, SNMP).
- Observability & Data Stack: Hands-on experience with Grafana, Prometheus/InfluxDB, ClickHouse, MySQL, and building custom metric exporters.
- Infrastructure as Code: Proficiency with modern configuration management tools (SaltStack, Ansible, Terraform) and GitHub CI/CD workflows.
- AI Integration: Daily, practical experience using LLM-based coding assistants and AI analytics tools in a professional software development workflow.
Benefits:
- Premium health, dental, and vision coverage with top-tier benefits.
- Generous performance-based bonus structures and long-term incentive programs.
- Daily provided meals and high-end office amenities in a premier, collaborative workplace.
- Fully funded access to top-tier developer tools, hardware, and AI platforms.
- Comprehensive continuous learning and conference allowance.
Salary:
- $175,000 – $235,000 + Performance-Based Bonus (Commensurate with experience)