Skip to main content

AI Infrastructure System Engineer

JobgetherIndia | AsiaToday
RemoteRustCloud & InfrastructurePythonKubernetesTerraformAnsibleLinuxFirmware

Job Description

AI Infrastructure System Engineer

India
Security & IT – IT /
Full-time /
Remote

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Infrastructure System Engineer  based in India.

This role offers the opportunity to build and operate large-scale infrastructure supporting advanced AI training and inference workloads.
You’ll engineer systems that manage thousands of GPUs with a strong focus on automation, reliability, performance, and availability.
The position goes beyond traditional infrastructure operations, emphasizing software-driven solutions and autonomous systems.
You’ll develop platforms that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal manual intervention.
Your work will span hardware, networking, storage, distributed systems, and AI workloads.
You’ll collaborate closely with infrastructure, hardware, networking, platform, and AI engineering teams to solve complex systems challenges.
It’s an ideal environment for an automation-focused engineer who enjoys building intelligent infrastructure at significant scale.

Accountabilities:

  • Design and build fleet automation systems capable of provisioning, validating, deploying, upgrading, repairing, and retiring GPU clusters with minimal human intervention.
  • Develop AI infrastructure agents that automate deployment workflows, investigate root causes, triage incidents, and support autonomous remediation.
  • Build fleet intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance.
  • Develop predictive capabilities that identify potential infrastructure failures before they affect customers or workloads.
  • Build software and automation systems that maximize GPU availability, utilization, performance, and reliability across large accelerator fleets.
  • Create automated validation frameworks for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage systems, and distributed AI workloads.
  • Develop internal infrastructure platforms and developer tools that enable infrastructure to be managed programmatically rather than through manual operations.
  • Continuously improve deployment velocity, system reliability, and operational efficiency through automation and software engineering.
  • Collaborate with hardware, networking, platform, and AI teams to identify infrastructure challenges and develop scalable solutions.
  • Apply strong systems thinking to problems spanning hardware and software components across large-scale AI infrastructure.

Requirements:

  • 3+ years of experience building distributed systems, infrastructure platforms, or large-scale backend software.
  • Strong software engineering skills in Python, Go, or Rust.
  • Proven experience developing platforms, automation systems, developer infrastructure, or similar software-driven infrastructure solutions.
  • Experience working with Linux and modern infrastructure technologies such as Kubernetes, Terraform, Ansible, or comparable tools.
  • Strong understanding of distributed systems and the ability to reason across hardware and software layers.
  • Passion for solving complex infrastructure problems through software and automation.
  • Strong automation-first mindset, with an instinct to build systems that eliminate repetitive manual tasks.
  • Ability to work effectively on complex technical challenges involving reliability, performance, scalability, and operational efficiency.
  • Strong collaboration skills and the ability to partner effectively with engineering teams across infrastructure, hardware, networking, and AI.
  • Experience with GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch, or related technologies is a plus.
  • Experience with InfiniBand or RoCE networking is advantageous.
  • Familiarity with bare-metal provisioning and infrastructure lifecycle management is desirable.
  • Experience supporting large-scale AI training or inference clusters is a plus.
  • Knowledge of hardware health monitoring and predictive failure detection is beneficial.
  • Experience with distributed storage systems is advantageous.
  • Familiarity with AI agents or autonomous infrastructure operations is a plus.

Benefits:

  • Opportunity to work on large-scale AI infrastructure supporting advanced training and inference workloads.
  • Exposure to complex systems spanning GPUs, networking, storage, distributed computing, and AI workloads.
  • Opportunity to build highly automated infrastructure systems and developer platforms.
  • Collaboration with multidisciplinary engineering teams working across hardware, networking, platform, and AI.
  • Environment focused on software-driven infrastructure, automation, scalability, reliability, and performance.
  • Opportunity to contribute to systems operating at significant GPU scale.
  • Remote-friendly job listing based in Bangalore, India.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
 
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
 
 
#LI-CL1
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
  • Design and build fleet automation systems capable of provisioning, validating, deploying, upgrading, repairing, and retiring GPU clusters with minimal human intervention.
  • Develop AI infrastructure agents that automate deployment workflows, investigate root causes, triage incidents, and support autonomous remediation.
  • Build fleet intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance.
  • Develop predictive capabilities that identify potential infrastructure failures before they affect customers or workloads.
  • Build software and automation systems that maximize GPU availability, utilization, performance, and reliability across large accelerator fleets.
  • Create automated validation frameworks for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage systems, and distributed AI workloads.
  • Develop internal infrastructure platforms and developer tools that enable infrastructure to be managed programmatically rather than through manual operations.
  • Continuously improve deployment velocity, system reliability, and operational efficiency through automation and software engineering.
  • Collaborate with hardware, networking, platform, and AI teams to identify infrastructure challenges and develop scalable solutions.
  • Apply strong systems thinking to problems spanning hardware and software components across large-scale AI infrastructure.

Requirements:

  • 3+ years of experience building distributed systems, infrastructure platforms, or large-scale backend software.
  • Strong software engineering skills in Python, Go, or Rust.
  • Proven experience developing platforms, automation systems, developer infrastructure, or similar software-driven infrastructure solutions.
  • Experience working with Linux and modern infrastructure technologies such as Kubernetes, Terraform, Ansible, or comparable tools.
  • Strong understanding of distributed systems and the ability to reason across hardware and software layers.
  • Passion for solving complex infrastructure problems through software and automation.
  • Strong automation-first mindset, with an instinct to build systems that eliminate repetitive manual tasks.
  • Ability to work effectively on complex technical challenges involving reliability, performance, scalability, and operational efficiency.
  • Strong collaboration skills and the ability to partner effectively with engineering teams across infrastructure, hardware, networking, and AI.
  • Experience with GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch, or related technologies is a plus.
  • Experience with InfiniBand or RoCE networking is advantageous.
  • Familiarity with bare-metal provisioning and infrastructure lifecycle management is desirable.
  • Experience supporting large-scale AI training or inference clusters is a plus.
  • Knowledge of hardware health monitoring and predictive failure detection is beneficial.
  • Experience with distributed storage systems is advantageous.
  • Familiarity with AI agents or autonomous infrastructure operations is a plus.

Benefits:

  • Opportunity to work on large-scale AI infrastructure supporting advanced training and inference workloads.
  • Exposure to complex systems spanning GPUs, networking, storage, distributed computing, and AI workloads.
  • Opportunity to build highly automated infrastructure systems and developer platforms.
  • Collaboration with multidisciplinary engineering teams working across hardware, networking, platform, and AI.
  • Environment focused on software-driven infrastructure, automation, scalability, reliability, and performance.
  • Opportunity to contribute to systems operating at significant GPU scale.
  • Remote-friendly job listing based in Bangalore, India.
The Rusty Bucket
Weekly curated Rust jobs delivered to your inbox.