Mirantis logo

Senior Data Platform Engineer

Mirantis
Remote
Posted 1 day ago

AI summary

Design and operate a production event-streaming platform for the k0rdent-ai platform, managing data movement and automation within an agile engineering environment.

Eligible from: Worldwide

Job Description

We are looking for an experienced data platform engineer to own the event-streaming backbone behind the k0rdent-ai platform — the multi-tenant control plane for enterprise GPU infrastructure. Every cluster provisioned, every GPU-hour consumed, and every tenant action produces events that must move reliably from our services into analytics, billing, observability, and downstream systems.  You will design and operate that pipeline end to end: the streaming platform itself, the connectors and change-data-capture flows that feed it, the schemas that keep producers and consumers compatible, and the automation that lets product teams onboard themselves without waiting on you. Working within an agile framework alongside engineering teams across the US, Europe, and India, you will make event data a dependable product surface rather than a best-effort side channel.

Main Responsibilities:

  • Design, deploy, and operate the production streaming platform on Kubernetes — cluster lifecycle, topics, partitioning and retention strategy, capacity planning, and upgrades.
  • Build and maintain change-data-capture and connector pipelines moving data between operational databases, the streaming layer, and analytical stores.
  • Own schema governance — a schema registry, compatibility rules, and versioning discipline so producer changes never silently break consumers.
  • Develop custom connectors, transformations, and stream-processing logic in Java or Go where off-the-shelf components fall short.
  • Build self-service onboarding so product teams can provision topics, schemas, and access through infrastructure as code rather than tickets.
  • Automate the platform with Terraform and GitOps — no manually configured clusters, no undocumented topic definitions.
  • Instrument the pipeline for production: metrics, lag and throughput monitoring, alerting, and dashboards that both the platform team and application owners can act on.
  • Enforce security and multi-tenant isolation across the streaming layer — mTLS, authentication, topic-level RBAC, encryption, and audit trails.
  • Guarantee delivery semantics that the business can rely on — ordering, idempotency, exactly-once where required, replay, and disaster recovery.
  • Partner with backend and SRE teams on event schema design, producer patterns, and failure modes; participate in on-call for the platform you own.
  • Mentor engineers on streaming architecture and raise the team's bar through design review and documentation.

Requirements

Required Skills/Abilities: 

  •  10+ years in software, data, or platform engineering, including deep hands-on ownership of a production event-streaming platform.
  • Expert-level Apache Kafka (or Confluent Platform/Cloud) — administration, tuning, troubleshooting, and capacity management at production scale.
  • Strong Kafka Connect and CDC experience — Debezium, JDBC, and custom connector or transformation development.
  • Proven schema-registry practice and compatibility management across many independent producers and consumers.
  • Solid programming ability in Java, Python, or Go — enough to build connectors, transformations, and platform tooling, not only configure them.
  • Strong SQL and relational database knowledge (PostgreSQL, MySQL, or equivalent), including replication and CDC mechanics.
  • Kubernetes fluency — deploying, operating, and debugging stateful workloads.
  • Infrastructure as code discipline (Terraform or equivalent) and CI/CD pipeline ownership.
  • Clear written English for design docs, runbooks, and asynchronous review across global time zones.

Must Have

  • Streaming: Apache Kafka or Confluent, Kafka Connect, Schema Registry, and stream processing.
  • Data Movement: CDC tooling (Debezium or equivalent), connector development, and pipelines into analytical stores.
  • Kubernetes Native: Kubernetes, Docker, and Helm-packaged workloads in production.
  • Automation: Terraform, CI/CD (Jenkins, GitHub Actions, or equivalent), and GitOps practice.
  • Data Stores: PostgreSQL at scale, plus at least one cloud data warehouse or search store (Snowflake, BigQuery, Elasticsearch, or equivalent).
  • Cloud: AWS or GCP — networking, IAM, managed Kubernetes.
  • Security & Observability: mTLS, OIDC/SSO, RBAC, and Prometheus/Grafana monitoring.

Nice to Have

  • Confluent certification, or contributions to Kafka-ecosystem open source.
  • Stream processing with Flink, Kafka Streams, or ksqlDB.
  • Event-driven architecture patterns — outbox, saga, event sourcing, CloudEvents.
  • Exposure to GPU infrastructure, AI/ML data pipelines, or telemetry at high cardinality.
  • Usage metering, billing, or chargeback pipelines built on event data.
  • Alternative brokers (NATS JetStream, Pulsar) and a view on the tradeoffs.
  • Go, for working directly in our backend services' producer code.
  • Data quality, lineage, or catalog tooling.
  • Compliance exposure — SOC 2, ISO 27001, or similar audit support.

Education and Experience:

  • Bachelor’s degree in Computer Science & Engineering or related field or 10 years related experience.

Ready to Apply?

Take the next step in your career journey

Apply Now

About the job

Posted on
Sep 8, 2026
Job type
Full-time
Location
Remote, inRemote

Keep looking

Related roles you might like

Explore more

Browse more jobs like this

Work arrangement

Disclaimer: Real Jobs From Anywhere is an independent platform dedicated to providing information about job openings. We are not affiliated with, nor do we represent, any company, agency, or agent mentioned in the job listings. Please refer to our Terms of Services for further details.