Back to Jobs
Lilt
AI & Machine Learning 1h ago

Freelance Multilingual Software Engineer for Terminal-Bench Task Evaluation

Lilt
TaiwanTaiwan
Freelance
Not Disclosed
Senior-Level

Job Description

Key Skills Required

Master these to land this role

Prompt Engineering17mFree Trial ✨
Start 10-Day Free Trial
Machine Learning41mFree Trial ✨
Start 10-Day Free Trial
Python ScriptingTerminal/CLI-based developmentNLP

Want to know if you're a match for this job?

Calculate My Match Score

We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.

We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches.

What You’ll Deliver

  • Task Engineering: Evaluating Coding Agents.

  • Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling.

  • Prompting & Translation: Identify failure points where AI does not work, in your native language.

  • Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary).

  • Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus).

  • Quality Assurance: Participate in a rigorous, 4-layer human quality control process (creation, human review, calibration review, and audit) alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity.

Qualifications

  • Experience: 5+ years of industry experience in software engineering.

  • Background: Proven track record at leading technology companies and/or graduation from top-tier engineering universities.

  • Language: Native or near-native fluency, with a deep understanding of its grammar, register, and phrasing rules. High English proficiency.

  • Technical Stack: Strong proficiency in Python, standard shell scripting, and data processing.

  • Workflow: Extensive experience with Terminal/CLI-based development workflows and a working familiarity with coding agents.

  • Domain Expertise: Deep technical understanding of multilingual text processing pitfalls, including:

    • Encoding/decoding robustness and Unicode normalization.

    • Locale-dependent conventions (collation, casing, non-Gregorian dates).

    • Text I/O, toolchain interoperability, and safe string operations.

    • (For specific languages) Bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts.

How would you rate this job post?

See what other professionals think about this role.

banner

Lilt is a highly disruptive, enterprise-grade AI translation and localization platform fundamentally designed to make the world's information universally accessible. Founded in 2015 by former Google Translate researchers Spence Green and John DeNero, and headquartered in San Francisco, California, the company operates as the ultimate contextual AI engine for global communications. Under the hood, Lilt seamlessly blends cutting-edge Large Language Models (LLMs) with a proprietary "Human Intelligence Layer"—empowering organizations to orchestrate complex translation workflows, automate technical documentation, and launch multilingual brand campaigns with unprecedented speed. Their primary target audience spans massive Fortune 500 enterprises (including Intel, Canva, and Lenovo), frontier AI developers, and high-security public sector agencies like the U.S. Department of Defense who desperately need defense-grade, SOC 2 compliant localization without sacrificing brand accuracy. What sets Lilt apart in the crowded localization ecosystem is its continuous model improvement and Agentic AI capabilities; instantly retraining its custom algorithms with every human translation to deliver massive cost reductions and up to 5x faster workflows, all while maintaining 100% data control across on-premise and air-gapped environments.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More