Back to Jobs
Pragmatike
AI & Machine Learning 16h ago

AI Infrastructure Engineer

Pragmatike
UkraineUkraine
🌍Czechia
BulgariaBulgaria
LatviaLatvia
SpainSpain
HungaryHungary
AlbaniaAlbania
LithuaniaLithuania
GreeceGreece
🌍Bosnia and Herzegovina
CroatiaCroatia
EstoniaEstonia
SerbiaSerbia
DubaiDubai
PolandPoland
ArmeniaArmenia
PortugalPortugal
ItalyItaly
MaltaMalta
TurkeyTurkey
MontenegroMontenegro
RomaniaRomania
Full-time
Not Disclosed
Senior-Level

Job Description

Key Skills Required

Master these to land this role

DevOpsBestseller 🔥
Learn in 63 Hours
Machine LearningBestseller 🔥
Learn in 42 Hours
Python ScriptingAI EngineerTensorFlow

Want to know if you're a match for this job?

Calculate My Match Score

About the Role

Pragmatike is recruiting on behalf of a fast-scaling, well-funded distributed cloud infrastructure startup building next-generation AI-native cloud services. The company is redefining how compute is delivered by providing GPU-powered infrastructure for AI/ML workloads, secure storage, and high-speed data transfer through a decentralized architecture that significantly reduces environmental impact compared to traditional cloud providers.

We are seeking a AI Infrastructure Engineer with strong experience in production-grade model serving and infrastructure for AI systems. This is a highly technical, hands-on role focused on building scalable, reliable, and efficient ML inference platforms powering real-time AI applications.

You will be responsible for designing and operating the core infrastructure that serves machine learning models at scale. You will work closely with infrastructure, platform, and applied AI teams to ensure high availability, low latency, and cost-efficient inference systems. Strong ownership, production mindset, and experience with distributed GPU systems are essential.

Your Responsibilities

  • Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent
  • Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models
  • Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers
  • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance
  • Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health
  • Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments
  • Own the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities
  • Define engineering best practices and contribute to platform scalability in a fast-moving startup environment

Required Qualifications

  • 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems
  • Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent
  • Strong background in container orchestration and operating GPU-based workloads in production
  • Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines
  • Proficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar)
  • Strong understanding of distributed systems, performance tuning, and production reliability engineering
  • Ability to effectively use AI coding assistants to accelerate development and debugging workflows
  • Ownership mindset with the ability to operate independently in a remote-first environment

Preferred Qualifications

  • Experience with ML platforms such as Kubeflow, MLflow, or KubeAI
  • Knowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference systems
  • Experience with cost optimization across different GPU types and inference workloads
  • Background in early-stage startups or greenfield infrastructure projects
  • Proven experience building production systems from scratch rather than maintaining legacy platforms

Why Join Us

  • Take ownership of critical infrastructure powering a rapidly scaling AI-native cloud platform
  • Build foundational ML inference systems from the ground up in a high-growth, well-funded startup
  • Work at the intersection of distributed systems, GPU computing, and sustainable cloud architecture
  • Gain deep expertise in next-generation AI infrastructure and large-scale model serving systems
  • Influence core engineering decisions and define best practices that will scale with the company.

How would you rate this job post?

See what other professionals think about this role.

banner

Pragmatike is a forward-thinking company that embodies the principles of pragmatism in its approach to innovation and problem-solving. With a name that reflects a practical and sensible outlook, Pragmatike likely operates in a cutting-edge industry where adaptability and efficiency are paramount. The company's emphasis on pragmatism suggests that it values effectiveness and results-driven strategies, making it a formidable presence in its sector. Pragmatike's mission is to provide solutions that are not only innovative but also grounded in reality, ensuring that its products or services are tailored to meet the actual needs of its customers. By combining creativity with a down-to-earth approach, Pragmatike aims to make a meaningful impact in the lives of its clients and contributors alike. The company's commitment to pragmatism also implies a focus on continuous learning and improvement, as it seeks to refine its methods and stay ahead of the curve in an ever-evolving landscape. With a strong foundation in practicality, Pragmatike is poised to achieve significant milestones and establish itself as a leader in its field. By fostering a culture of innovation and collaboration, Pragmatike encourages its team members to think critically and develop novel solutions that are both effective and sustainable. Through its dedication to pragmatism, Pragmatike strives to create value for all stakeholders, from customers and partners to employees and the wider community. As the company continues to grow and expand its operations, it remains steadfast in its pursuit of excellence and its mission to make a lasting difference.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More