Quality Assurance Engineer (AI & Automation)
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
What you’ll do
Own quality strategy end-to-end: define, implement, and continuously evolve testing standards across functional, non-functional, and AI-specific dimensions, ensuring quality is embedded from requirements through production.
Build and maintain non-functional test automation: design and run performance, load, and stress test suites (k6, JMeter, Gatling etc.) integrated directly into CI/CD pipelines, with quality gates that protect every release.
Design and operate (or contribute to) LLM/AI eval frameworks: establish evaluation pipelines (using tools such as DeepEval, Langfuse etc.) to assess AI feature quality across metrics including accuracy, hallucination rate, relevance, faithfulness, and safety.
Test AI features and agentic behaviours: validate non-deterministic outputs, prompt variability, model regression, guardrail enforcement, and multi-step agent task-completion rates as first-class quality concerns.
Champion shift-left and continuous testing: embed QA into planning, design review, and sprint ceremonies so defects are caught before they're coded, not after they ship.
Drive a quality engineering culture: act as a quality advocate across engineering, product, and AI teams; run blameless post-mortems, define quality metrics, and make test coverage and reliability visible to the whole organization.
Accelerate delivery through AI-assisted tooling: use AI coding assistants, self-healing automation, and intelligent test prioritization to increase the leverage of every hour spent on quality work.
Build observability into production: define and monitor post-release quality signals, model drift indicators, and SLO thresholds so the team can distinguish a regression from expected non-determinism.
What you bring
Must-Haves
Traditional QA foundations: solid understanding of deterministic testing: test planning, test case design, functional/regression/exploratory testing, defect lifecycle management, and quality metrics.
Test automation engineering: deep expertise in writing and maintaining automated test suites using modern frameworks (Playwright, Cypress, or similar) with at least one modern programming language, such as TypeScript (strongly preferred) or Python, specifically for building robust test libraries.
Non-functional test automation: hands-on experience designing and running performance, load, and stress tests with tools such as k6 or JMeter, including CI/CD integration and threshold-based quality gates.
AI/LLM testing literacy: practical understanding of what makes AI systems non-deterministic, and experience (or strong working knowledge) of testing LLM-based features for hallucination, consistency, safety, and latency.
Eval framework awareness: a working understanding of LLM evaluation concepts: scoring metrics (BLEU, ROUGE), LLM-as-judge patterns, and familiarity with at least one eval framework (DeepEval, RAGAS, etc.).
CI/CD and continuous testing: experience integrating test suites into pipelines (GitHub Actions, CircleCI, or equivalent) with a shift-left mindset that treats test failures as blocking signals, not background noise.
Quality ownership mentality: demonstrated ability to own quality outcomes, not just execute tasks; comfort setting standards, raising risk flags, and influencing cross-functional teams.
Solid understanding of architectural patterns, microservices, and API testing (REST, gRPC) using tools like Postman or custom frameworks.
Experience with containerization technologies (Docker) and orchestration (Kubernetes) as they relate to scalable testing environments.
Nice-to-Haves
Experience with agentic systems testing: validating goal-completion rates, guardrail enforcement, and multi-step reasoning chains in LLM agent workflows.
Familiarity with AIOps/MLOps/LLMOps concepts: prompt versioning, model monitoring, canary deployments, and drift detection.
Experience with accessibility or security testing as part of a broader non-functional quality practice.
Background in red teaming, adversarial input testing, or prompt injection validation.
Experience working in a startup or scale-up environment where processes are built from scratch rather than inherited.
Why Togal?
Join a dynamic team of AI-native engineering team.
AI-native from day one - you'll be building the quality discipline for a product that uses AI at its core, making every quality decision novel and impactful.
Be the quality voice, not a quality follower - this role has direct influence over how Togal defines and measures product excellence.
A culture of innovation, continuous learning, and high growth.
Comprehensive benefits
Competitive compensation package and flexible work arrangements.
We are an equal opportunity employer committed to building a diverse team. We welcome applications from candidates of all backgrounds who are passionate about using technology to transform the construction industry.
Join us in revolutionizing pre-construction estimating with the power of AI!
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
Product Manager - Autonomous Driving (Mapping & Localization)
Torc Robotics
United StatesLead Product Manager, Platform and Integrations
Lirio
United StatesSolutions Architect, Inbound AI Deployments
Hippocratic AI
United StatesVice President, SE - Majors
Snowflake
United StatesTogal.AI
View Company ProfileTogal.AI (operating at togal.ai) is an AI-powered takeoff tool built by estimators that automatically detects, measures, and compares directly from your drawings — with up to 98% accuracy. Founded in 2019 and headquartered in Miami, Florida, Togal.AI helps construction professionals complete faster, more accurate AI-assisted takeoffs — without replacing the estimator. Under the hood, Togal.AI uses proprietary AI algorithms to automatically and accurately detect, label, and measure project spaces & objects within seconds. This allows construction teams to save time and increase accuracy. Backed by $5 million in a pre-Series A SAFE round with a $50 million valuation cap.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.
