Staff / Principal MLOps Engineer
United StatesJob Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Your Role
We're seeking a Staff / Principal MLOps Engineer to join our team. Our ML footprint has grown quickly alongside the business: batch models that process records in the backend, real-time models that serve recommendations to job seekers, and daily pipelines that process every available job across the US and Canada. The layer we have not yet built is the observability and traceability around all of it. Today, when a model regresses, a job fails, or a recommendation looks wrong, especially where LLMs are involved, tracing the cause and reproducing it takes far longer than it should. You will own that problem: assess our ML pipelines and data architecture with clear eyes, decide what to build and in what order, and then build it. This is a greenfield mandate, influencing production models and users. We are open to running this as a six-month contract or as a full-time hire, depending on fit and what you are looking for.
What You'll Own
Assessment and plan: Evaluate our current pipelines, data architecture, and ML workflows, and produce a prioritized, opinionated plan for what needs to change and why.
AI/ML observability: Architect our AI/ML observability and traceability from the ground up: model and data monitoring, regression detection, lineage, and the ability to reproduce a questionable recommendation on demand, including for LLM-based systems.
Systems design: Design data and ML systems that are anchored in customer needs and built to last, with clear tradeoffs documented so the team can build on them.
Implementation: Rebuild and harden pipelines, upgrade the data architecture, and ship the improvements.
Reliability and standards: Raise the bar on reliability and data quality, establishing the patterns and practices the rest of the team can run with.
Dependency and security hygiene: Keep the stack current and secure: framework and package upgrades across services and model images, and vulnerability remediation carried out without destabilizing production.
Required Experience
Staff or principal-level experience in MLOps, ML platform, or ML infrastructure
Experience standing up MLOps practice: CI/CD for models, experiment tracking, feature stores, and model monitoring
Experience building AI/ML observability and traceability in production: detecting regressions, diagnosing failures, tracing a prediction back to the inputs that produced it, and reproducing issues after the fact
Experience operating models in both batch and real-time serving contexts, with an understanding of how the reliability requirements differ
A track record of walking into complex, fast-grown systems, diagnosing the real problems, and materially improving them
Strong systems design ability: you can translate customer and product needs into durable, scalable architecture, write it down clearly, and stay close enough to the code to implement it yourself
Depth in classical ML methods in production, plus practical exposure to LLM-based systems and what it takes to observe and evaluate them once they are live
Deep experience building and operating production data pipelines and ML workflows at scale
Fluency across the modern ML and cloud stack: orchestration, containerization, infrastructure as code, CI/CD, model serving, and monitoring
Bonus Points
Experience evaluating and working with AI/ML observability or LLM evaluation vendors, including clear judgment on when to buy and when to build
Background in mission-driven, workforce, or government-adjacent data environments
Publications, presentations, blog posts, or other public artifacts showcasing your expertise and knowledge of best practices in MLOps
Comfort mentoring and leveling up a small data and engineering team while you build
Our Tech Stack for Data
Languages: SQL, Python
Data orchestration and transformation: Airflow, dbt
Data storage and warehousing: PostgreSQL, Redshift, MongoDB
Machine learning and model serving: AWS SageMaker (PyTorch models, artifact upload to S3, model registration), serving real-time and batch inference
Visualization and reporting: Looker, Quicksight
Infrastructure: AWS (S3, Redshift), GitHub Actions for CI/CD
Your Education
Your alma mater isn't our focus. Your grit, hunger, and drive are. If you learn continuously, tackle challenges head-on, and know your strengths and gaps intimately, you're our person.
Location
[CA/US Remote] We are open to candidates living anywhere in Canada or the US. For candidates living in Toronto, our office is conveniently located at 325 Front St West (a short walk from Union Station). You are welcome to come in on a hybrid schedule.
Travel Expectations
Although this role is remote, you may be expected to travel up to once per quarter for off-sites and team gatherings.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at FutureFit AI
Explore Top Companies in this Space
Learnlight
EdTech / Corporate Training / B2B SaaS
GenPeach AI
Artificial Intelligence / Generative AI / Enterprise Software
HireDigital
Human Resources, Recruitment Technology

HR POD
Human Resources
FutureFit AI
View Company ProfileFutureFit AI (operating via futurefit.ai) is a premier AI-powered workforce technology enterprise engineered to build employment pathways and bridge the gap between talent, training, and employment at scale. Founded in 2018 and headquartered in New York City with significant remote operations, the company fundamentally transforms how individuals navigate career transitions in the face of rapid automation and labor market disruptions. Moving far beyond traditional, one-size-fits-all career advice or static job boards, FutureFit AI natively unifies comprehensive labor market intelligence (leveraging over 350 million talent profiles and 50 million job postings), intelligent skills mapping, and personalized career roadmaps into a single, cohesive ecosystem. The platform empowers job seekers, employers, and workforce development organizations to align skills with in-demand opportunities, delivering tailored guidance from onboarding to advancement. Under the hood, their sophisticated Career Copilot and Pathways Platform utilize proprietary machine learning algorithms designed to eliminate systemic biases and recommend actionable upskilling and reskilling resources. Backed by investors like JPMorgan Chase and Acumen America, and trusted by leading enterprise and government partners, FutureFit AI remains a definitive cornerstone of the modern workforce development landscape, ensuring that no one is left behind in the future of work.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.