Back to Jobs
MeridianLink
AI & Machine Learning 2h ago

AI Engineer – Trust & Explainability (AI Platform)

MeridianLink
United StatesUnited States
Full-time
Not Disclosed
Mid-Level

Job Description

Key Skills Required

Master these to land this role

Python2h 41mFree Trial ✨
Start 10-Day Free Trial
LLM observability and tracingLLM evaluation and testingOpenTelemetryAgent frameworks and multi-agent orchestration

Want to know if you're a match for this job?

Calculate My Match Score

The AI Platform team is building the shared runtime for every AI agent at MeridianLink, including model gateway, orchestration, memory, and observability, delivered as reusable services. This role focuses on the trust and explainability layer: tracing, evaluation, and explanation capabilities that allow engineers to understand agent workflows and enable product teams to show end customers the right level of explanation to trust an agent's output.

The role has a strong dotted-line relationship with MeridianLink's Security Operations team, ensuring platform trust guarantees and security threat models are aligned.

About the Opportunity

This is a deeply technical, hands-on engineering role integral to the platform itself. Responsibilities include:

  • Writing code to trace agent actions across multi-agent workflows
  • Adopting and extending leading open-source explainability and observability frameworks
  • Building primitives for product teams to surface explanations to lenders, credit unions, and borrowers

Candidates should have real hands-on experience with LLM-based systems and a strong interest in trust, tracing, and evaluation. Expertise is not required upfront; this role offers growth alongside a Staff engineer team lead and architects.

What it means to be an L2 AI Engineer at MeridianLink

L2 AI Engineers handle a broad range of work independently, from bug fixes to feature development. They specialize in their domain and deepen expertise while writing code trusted by other engineers. They actively use AI-assisted development tools.

Technical Execution & Delivery

  • Completes features and bug fixes independently with minimal guidance
  • Designs components within a well-defined scope and escalates complex design questions
  • Participates in code reviews and provides constructive feedback
  • Proactively surfaces blockers rather than waiting for check-ins

Craft & Professionalism

  • Writes comprehensive tests covering shipped functionality
  • Monitors and addresses issues with their own work
  • Documents decisions and implementation details for future reference

Agent Tracing & Explainability

  • Instruments agents to trace conversations from user messages through every model call, tool call, retrieval, handoff, and decision to the final response
  • Builds correlation to link actions across multiple agents into a single readable trace
  • Transforms raw trace data into human-readable explanations, detailing what the agent did, relied on, and why it chose that path

Customer-Facing Trust & Explanation

  • Understands the distinction between explanations for engineers and those for borrowers or loan officers, building for both audiences
  • Collaborates with product engineers to determine what agents should disclose at the product surface, including actions, data used, confidence levels, and verification steps
  • Applies judgment to determine the appropriate level of explanation for regulated lending and account-opening contexts

Trust & Explainability Frameworks

  • Evaluates open-source frameworks for LLM observability, tracing, and evaluation, aligning them with platform needs
  • Integrates and extends existing frameworks rather than rebuilding them, filling gaps where necessary
  • Ensures platform instrumentation aligns with emerging standards to maintain trace portability across tools

Agent Evaluation & Quality

  • Measures non-deterministic output quality using golden datasets, rubric and LLM-as-judge scoring, and regression baselines
  • Builds evaluation checks that run in CI to test model swaps, prompt changes, and tool changes before customer exposure
  • Treats evaluation data as versioned, reviewed code

Safety & Tenant Isolation

  • Identifies common attack patterns against LLM applications (prompt injection, jailbreaks, tool misuse, data exfiltration) and writes tests for them
  • Understands tenant isolation as a platform guarantee, building tests to ensure one customer's agent cannot access another's data
  • Contributes to shared guardrail layers to protect every agent without re-implementing protections

What Success Looks Like

In the first few months, a successful hire will have shipped tracing for a full agent workflow, readable by engineers unfamiliar with the codebase, and contributed working components to the platform's evaluation harness. Over the first year, success includes:

  • A first version of the customer-facing explanation primitive live in at least one product
  • Automated evaluation and isolation checks running on every agent release
  • Platform providing clear answers when engineers or customers ask about agent actions or trustworthiness

Engineers who thrive here enjoy the intersection of AI and rigor, aim to become the team's go-to on agent tracing and explainability, and are motivated by making AI transparent and trustworthy.

Key Responsibilities

Multi-Agent Tracing & Explainability

  • Build tracing across the platform's gateway, orchestration, memory, and tool layers, following defined designs
  • Correlate actions across agents in a single multi-agent workflow, including handoffs, parallel branches, and retries
  • Create developer-facing trace views readable by non-writers of the agent
  • Build explanation layers that turn trace data into human-readable accounts of agent decisions

Customer-Facing Trust

  • Build platform primitives for product teams to surface to end users, including explanation records, confidence and provenance metadata, and summaries of what agents relied on
  • Partner with product engineers on agents like Document Request Agent and MLM agents to integrate these primitives
  • Iterate on explanation formats based on feedback from product teams and customer-facing staff

Trust & Explainability Frameworks

  • Evaluate and integrate open-source observability, tracing, and evaluation frameworks into the platform runtime
  • Extend frameworks where needed and contribute fixes and extensions upstream
  • Build missing trust and explainability tooling not provided by the ecosystem

Evaluation Tooling

  • Build and maintain components of the platform's evaluation framework, including golden dataset management, test runners, scoring pipelines, and regression reporting
  • Develop tooling to help teams create and version golden datasets from de-identified real traffic and synthetic cases
  • Run model and prompt comparisons for shared platform components and report changes

Safety & Isolation Testing

  • Build red-team and adversarial test suites for prompt injection, jailbreaks, tool misuse, and data exfiltration, running them in the release process
  • Create automated test suites to prove agents on the shared runtime cannot cross tenant boundaries through memory, retrieval, tool calls, or model context
  • Share red-team findings and new attack patterns with Security Operations, integrating threat models into platform tests

Collaboration & Growing Others

  • Participate in design discussions and code reviews, giving and receiving constructive feedback
  • Support onboarding of L1 AI Engineer teammates, sharing context and helping them get unblocked
  • Contribute to documentation to reduce tribal knowledge on the team

Qualifications

Required Experience

  • 3+ years of professional software engineering experience, delivering features independently in production
  • Solid understanding of algorithms, data structures, and software design fundamentals
  • Proficiency with standard development tooling: Git, Docker, automated testing, and modern scripting languages
  • Active daily use of AI-assisted development tools
  • Bachelor's degree in Computer Science, Software Engineering, or equivalent experience
  • Hands-on experience building software integrating large language models (LLM APIs, agent frameworks, RAG pipelines, etc.) in production or substantial personal/open-source projects
  • Proficiency in Python with production experience; TypeScript is a plus
  • Hands-on experience in at least one of: LLM observability and tracing, LLM evaluation and testing, agent frameworks and multi-agent orchestration, or application security testing
  • Experience with distributed tracing or observability tooling in production systems
  • Strong automated testing instincts, including testing non-deterministic systems
  • Experience running workloads on Azure or AWS, including identity and access management, networking, and secrets management

Preferred Qualifications

  • Experience with OpenTelemetry, including GenAI semantic conventions, or OpenLLMetry
  • Experience with LLM observability and evaluation tools (Langfuse, Arize Phoenix, LangSmith, Braintrust, promptfoo, DeepEval, or equivalent)
  • Contributions to open-source AI observability, evaluation, or agent framework projects
  • Experience with cloud-managed model services such as AWS Bedrock or Azure OpenAI
  • Experience building or operating multi-tenant SaaS systems with tenant isolation requirements
  • Prior experience in financial services, fintech, or regulated industries where explainability shaped technical decisions
  • Experience building developer-facing debugging or visualization tools

How would you rate this job post?

See what other professionals think about this role.

banner

MeridianLink is a massive, publicly traded FinTech and enterprise software powerhouse fundamentally designed to digitize and automate the entire lending and account opening lifecycle for financial institutions. Founded in 1998 and headquartered in Costa Mesa, California, the company operates as the digital backbone for nearly 2,000 banks, credit unions, and consumer reporting agencies. Under the hood, MeridianLink seamlessly consolidates complex financial workflows—delivering highly robust Loan Origination Systems (LOS), digital account opening platforms, and advanced data verification solutions across consumer, mortgage, and indirect lending products. Their primary target audience spans progressive banking executives, loan officers, and financial IT leaders who desperately need to replace fragmented, legacy systems with a unified, 100% cloud-native ecosystem that accelerates approvals and maintains strict regulatory compliance. What sets MeridianLink apart in the fiercely competitive banking software landscape is its staggering multi-decade legacy of innovation and sheer processing scale; continuously leveraging AI-enabled decisioning and a massive partner integration network to securely power millions of loan applications and transform traditional lending into a fast, highly personalized, and frictionless digital experience.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More