AI summary
As a Software Engineer for Infrastructure, you'll design and implement services for a GPU-as-a-Service platform, focusing on building APIs and workflows for server provisioning and Kubernetes cluster management.
Job Description
We are looking for an experienced Software Engineer, Infrastructure to design and implement the Infrastructure Services that power our GPU-as-a-Service platform. You will build the control plane that turns high-level API calls into real infrastructure actions — enrolling bare-metal servers, provisioning them, and assembling them into multi-tenant Kubernetes clusters on high-performance hardware.
You will own the full lifecycle of the infrastructure-level services—from the Server and MachineType APIs down to the provisioning workflows and reconciliation loops that keep the platform's view of hardware consistent with physical reality.
Key Responsibilities
• Infrastructure API Design: Design, build, and maintain the versioned REST and gRPC APIs for bare-metal server lifecycle, MachineType definitions, and cluster CRUD operations.
• Provisioning & Workflow Development: Develop the asynchronous workflows that drive server enrollment, inspection, OS provisioning, and cluster bring-up, exposing durable status to callers.
• Implementation: Implement and maintain the consoles and interfaces that visualize hardware inventory, provisioning progress, and cluster health.
• System Reliability: Design the error-handling models, idempotency guarantees, and reconciliation loops necessary to manage long-running provisioning operations reliably.
Requirements
• API Development: Strong experience designing RESTful APIs or gRPC services. You understand API versioning and gateway patterns.
• Proficiency in Go (preferred for backend/Kubernetes ecosystem)
• Kubernetes Knowledge: Deep understanding of Kubernetes primitives and controller/reconciler patterns. You will be interacting with systems like k0rdent, Metal3, and Cluster API to translate high-level API calls into infrastructure actions.
• Bare-Metal Provisioning: Hands-on experience with bare-metal provisioning flows — BMC/Redfish, PXE/iPXE, image management, and hardware inspection.
• Asynchronous Systems: Experience building workflow-driven or event-driven systems (e.g., Temporal) where operations are long-running and state must remain consistent across retries and failures.
Preferred Qualifications
• State Reconciliation: Experience building informers or reconciliation bridges that keep an external datastore consistent with Kubernetes resource state.
• Multi-Tenancy: Experience building platforms where strict data and network isolation between tenants is required.
• Infrastructure-as-Code: Familiarity with Terraform/OpenTofu and GitOps-driven configuration (ArgoCD or Flux).
• Hardware Domain: Familiarity with GPU server hardware, DPUs/NICs, and high-performance datacenter fabrics.
About the job
- Posted on
- Aug 28, 2026
- Job type
- Full-time
- Location
- Remote, USA, usRemote
Keep looking
Related roles you might like
Director, Presales Solution Architecture - NeoCloud
Mirantis
Senior Software Systems Engineer (Storage) - remote in the US
Mirantis
Sales Development Representative (Remote on the West Coast)
Mirantis
Senior Software Engineer (Golang) - remote in the US
Mirantis
Senior Software Systems Engineer (Storage) - remote in the US
Mirantis
Enterprise Architect – Infrastructure Advisory _ remote in the US
Mirantis
Systems Architect, Enterprise Infrastructure _remote in the EU
Mirantis
Technical Product Marketing, Director - Remote (US)
Mirantis
Disclaimer: Real Jobs From Anywhere is an independent platform dedicated to providing information about job openings. We are not affiliated with, nor do we represent, any company, agency, or agent mentioned in the job listings. Please refer to our Terms of Services for further details.
