landing_page-logo
FluidStack logo

Senior Data Center Operations Engineer

FluidStackNew York City, New York

Automate your job search with Sonara.

Submit 10x as many applications with less effort than one manual application.1

Reclaim your time by letting our AI handle the grunt work of job searching.

We continuously scan millions of openings to find your top matches.

pay-wall

Job Description

About FluidStack

Fluidstack is the AI Cloud Platform. We build GPU supercomputers for top AI labs, governments, and enterprises. Our customers include Mistral, Poolside, Black Forest Labs, Meta, and more.

Our team is small, highly motivated, and focused on providing a world class supercomputing experience. We put our customers first in everything we do, working hard to not just win the sale, but to win repeated business and customer referrals.

We hold ourselves and each other to high standards. We expect you to care deeply about the work you do, the products you build, and the experience our customers have in every interaction with us.

You must work hard, take ownership from inception to delivery, and approach every problem with an open mind and a positive attitude. We value effectiveness, competence, and a growth mindset.

About the Role

Join the infrastructure team as a Senior Data Center Operations Engineer. You'll own critical infrastructure that powers AI innovation. This role demands extreme ownership and an automation-first mindset. You'll manage complex data center operations while building tooling that eliminates manual work. Your mission: ensure our GPU supercomputers run flawlessly 24/7 for the world's leading AI labs.

Focus

  • Own regional data center operations end-to-end. Manage power, cooling, and rack infrastructure across multiple sites.

  • Develop high-efficiency rack designs. Maximize space and power utilization to save millions in energy and infrastructure costs.

  • Build automation tools that eliminate routine tasks. Every manual process is an opportunity to code a solution.

  • Configure and manage global PDU infrastructure. Integrate with monitoring systems to calculate PUE, generate alerts, and create reports.

  • Lead DCIM implementation and adoption. Track assets across multiple locations, reduce asset retrieval time, and maintain 99.99%+ accuracy.

  • Drive automation initiatives. Integrate tracking systems with DCIM for real-time asset lifecycle updates.

  • Design reports and dashboards. Identify rack and power density improvements to drive efficiency and capacity optimization.

  • Maintain policy and procedure documents. Update SOPs and MOPs for compliance and efficiency.

  • Utilize ticketing and knowledge base applications. Leverage Jira, and Confluence to manage workflow and documentation.

  • Document everything. Write clear procedures that enable others to execute flawlessly.

About You

  • 5+ years managing data center operations at scale. Experience with hardware integration, capacity planning, and infrastructure optimization.

  • Proven track record achieving 99.99%+ accuracy in physical audits across multiple regions and countries.

  • Experience managing DCIM implementations. You've tracked 100K+ assets and reduced retrieval times by 75%.

  • Strong vendor management skills. Experience with ITAD relationships, hardware disposals, and generating revenue from e-waste.

  • Expertise in power infrastructure. Knowledge of PDU configuration, PUE calculations, and energy optimization.

  • Experience leading large-scale migrations. You've executed 10+ full-cage relocations maintaining continuous service.

  • Automation mindset. You've integrated RFID tracking with DCIM and improved data accuracy from 80% to 99.999%.

  • Excellent vendor management skills. You negotiate effectively and hold partners accountable.

  • Strong technical documentation skills. Experience creating and maintaining SOPs, MOPs, and training materials.

  • Data-driven approach. You create dashboards and reports that drive infrastructure decisions.

  • Extreme ownership mentality. You see problems through from identification to resolution.

Nice to have

  • Experience with GPU infrastructure and high-performance computing environments

  • Familiarity with AI/ML workloads and their infrastructure requirements

  • Knowledge of liquid cooling systems for high-density compute

  • Experience building custom monitoring and automation tools

  • Background in hyperscale or cloud data center operations

Benefits

  • Competitive total compensation package (salary + equity).

  • Retirement or pension plan, in line with local norms.

  • Health, dental, and vision insurance.

  • Generous PTO policy, in line with local norms.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Automate your job search with Sonara.

Submit 10x as many applications with less effort than one manual application.

pay-wall