Software Test Engineer (AI), TikTok Search

TikTok · Singapore

Sector
AI
Function
Product & Engineering
Level
Mid-Level
Employment type
Full Time
Posted
2026-08-05
Source
mycareersfuture

The TikTok Search Engineering Test Team owns the end-to-end quality of TikTok's global search experience—from query understanding, retrieval, and ranking to generative search experiences (LLM Answers and RAG) and search safety. TikTok Search operates across hundreds of markets and languages, connecting billions of users every day with the videos, creators, products, and information they are looking for. Every relevance regression, latency fluctuation, or unsafe result directly impacts user trust and business outcomes.We are redefining how search quality is built. By leveraging Large Language Models (LLMs), AI agents, and intelligent automation, we are reimagining the entire search testing lifecycle—from query set construction and relevance evaluation to ranking diagnostics, LLM-generated answer evaluation, RAG pipeline verification, and search safety red teaming.Our mission is to build next-generation AI-powered evaluation and testing capabilities that enable TikTok Search to iterate faster, ship smarter, and remain trustworthy at global scale. You will work at the intersection of Search Engineering, Quality Engineering, and Generative AI, helping shape the future of AI-native Search Quality at TikTok.As a Software Test Engineer (AI) on the Search team, you will:- Design and deliver AI-powered testing and evaluation solutions for TikTok Search, covering query understanding, retrieval, ranking, generative search (LLM Answers and RAG), personalization, and multimodal and vertical search scenarios such as video, image, e-commerce, user, and creator search.- Build LLM-based evaluation pipelines for search relevance, answer faithfulness, groundedness, freshness, coverage, diversity, and safety; construct and continuously maintain multilingual, multi-market query sets and golden datasets that reflect real user intent.- Drive change impact analysis across the search stack—including index, retrieval, ranking models, and prompt and model changes—through AI-assisted diff analysis, offline replay, side-by-side (SBS) evaluation, and interleaving experiments; strengthen release quality gates and pre-launch verification capabilities.- Partner closely with search engineering, algorithm, product, and content operations teams to identify quality and engineering productivity pain points across the search lifecycle, propose practical AI-powered solutions, and drive them into production.- Improve engineering productivity for the search organization by building AI agents and tooling for automated test case generation, log-driven bug analysis, root-cause attribution, and self-healing regression suites.QualificationsMinimum Qualifications- Bachelor's degree or above in Computer Science, Software Engineering, Information Retrieval, Data Science, or a related technical discipline.- Experience in software quality assurance or software testing, preferably for search, recommendation, ranking, content understanding, or other large-scale data- or algorithm-driven systems.- Solid understanding of software testing methodologies, debugging, and quality assurance best practices across the software development lifecycle, including offline evaluation, online A/B testing, and production monitoring.- Familiarity with distributed systems and search and recommendation infrastructure fundamentals (e.g. inverted indexes, vector indexes, retrieval, ranking, embeddings), plus basic knowledge of iOS and/or Android as search entry points.- Hands-on experience in one or more testing domains, including automation testing, performance and stress testing, algorithm evaluation, or test platform and tool development.- Strong interest in Generative AI and Large Language Models, with an understanding of common LLM capabilities, evaluation paradigms (e.g. LLM-as-a-Judge, RAG evaluation, and human evaluation), and real-world search applications.Preferred Qualifications- Proficiency in one or more programming languages such as Python, Go, Java, or C++.- Hands-on experience evaluating or testing search, recommendation, ranking, or LLM and RAG systems, including relevance labeling, SBS evaluation, NDCG, MRR, and Recall analysis, faithfulness and groundedness measurement, or hallucination detection.- Experience developing test automation frameworks, evaluation platforms, query set management tools, or engineering productivity tools for search and algorithm teams.- Hands-on experience with Generative AI–related projects, including LLM application development, AI agent evaluation, model evaluation, AI-assisted testing, prompt engineering, fine-tuning, or building AI testing frameworks.- Experience leveraging AI tools (e.g. LLMs, AI coding assistants, or AI agents) to improve testing efficiency, automate test case generation, bug analysis, or search quality workflows.- Experience with multilingual and cross-market quality work, content safety testing, or vertical search domains such as e-commerce, video, or creator search.

Apply on mycareersfuture →
AI Retrieval-Augmented Generation (RAG) Search Engine Technology Automated Tools organizational pain-points Test Case Generation Generative AI Application Development and Deployment E-Commerce