Senior GenAI Quality Engineer & Solution Analyst – LLM/RAG/Agentic Testing _ Contract

NTT Singapore · Singapore

Sector
AI
Function
Product & Engineering
Level
Junior
Employment type
Contract
Posted
2026-09-25
Source
mycareersfuture

Employment type: ContractDuration: 12 monthsExperience: 6–9 yearsCurrent work location: AlexandraExpected location from 2027: Punggol Digital District (PDD)Work arrangement: Primarily work from office, subject to the client’s prevailing policy.Job summaryWe are seeking a Senior GenAI Quality Engineer and Solution Analyst to analyse, test and validate enterprise Generative AI applications within a regulated financial-services environment.This is not a conventional manual testing role. The successful candidate must combine strong UI/API automation with hands-on testing of production LLM, RAG, conversational AI and agentic applications. The role also requires the ability to translate ambiguous business requirements into solution flows, acceptance criteria, evaluation scenarios and release evidence.Key responsibilitiesDefine end-to-end, risk-based test strategies covering web UI, REST APIs, backend services, databases, enterprise integrations and GenAI components.Validate LLM and RAG responses for task completion, grounding, relevance, factual accuracy, consistency, citations, hallucination risk and safe failure.Create evaluation datasets, golden test sets, quality thresholds and repeat-run test approaches for non-deterministic AI outputs.Test conversational and agentic behaviour, including multi-turn context, memory, tool selection, tool inputs and outputs, state transitions and termination conditions.Validate retries, timeouts, handoffs, human-approval checkpoints, fallback behaviour and recovery from partial failures.Perform functional, integration, regression, exploratory, negative, resilience, security and basic performance testing.Design API tests covering authentication, authorisation, contracts, validation, error handling, idempotency, rate limits and downstream failures.Develop and maintain UI and API automation using Playwright, Cypress, Selenium, PyTest, REST Assured, Postman or equivalent tools.Analyse logs, traces, network calls, payloads and database records to isolate application, model, retrieval, data, integration and platform defects.Validate prompt-injection resistance, restricted-data handling, access controls, auditability and safe responses.Translate business needs into user journeys, functional requirements, interface behaviours, decision rules, acceptance criteria and non-functional requirements.Map interactions across user interfaces, APIs, prompts, models, retrieval components, enterprise data sources, agent tools and downstream systems.Identify unclear requirements, missing controls, integration assumptions, failure scenarios and operational gaps before development begins.Produce practical artefacts including process flows, sequence diagrams, interface specifications, decision tables and traceability matrices.Assess solution trade-offs involving quality, cost, latency, security, maintainability and operational risk.Maintain traceability from business requirements through implementation, test scenarios, evaluation results and release evidence.Integrate stable, high-value automation into CI/CD pipelines.Provide evidence-based release recommendations covering known limitations, residual risks and production-monitoring requirements.Communicate defects and quality risks clearly to product owners, architects, engineers, security teams and business stakeholders.Support post-release monitoring and continuous improvement of GenAI quality.Mandatory requirements4 + years of hands-on software quality engineering, SDET or test-automation experience.Recent experience testing a real production or enterprise GenAI, LLM, RAG, chatbot or agentic application.Strong UI and REST API automation experience.Hands-on Playwright, Cypress or Selenium experience.Hands-on Postman, REST Assured, PyTest or equivalent API automation.Working knowledge of Python, Java, JavaScript or TypeScript.Experience validating hallucination, grounding, factuality, relevance, citations and multi-turn context.Experience testing retrieval quality, document ingestion, chunking, embeddings, reranking or vector-search behaviour.Experience with repeat-run evaluation, golden datasets and threshold-based acceptance.Experience analysing application logs, model traces, API payloads and database records.Strong requirements analysis, risk assessment, negative testing and traceability.Experience with Git, pull requests, CI/CD, test reporting and defect-management tools.Strong written and verbal stakeholder communication.Preferred experienceRagas, DeepEval, LangSmith, Langfuse, OpenTelemetry or equivalent.Agent tool-call, memory/state, HITL approval and failure-recovery testing.Banking, financial services, government or another regulated environment.Security, prompt-injection or adversarial GenAI testing.Kubernetes, OpenShift, AWS, Azure or containerised deployments.Accessibility, service virtualisation, contract testing or synthetic monitoring.Interested candidates are kindly requested to email their CV with their experience to [email protected] look forward to your application!

Apply on mycareersfuture →
AI Chatbot Retrieval-Augmented Generation (RAG) Selenium WebDriver Software Quality Assurance Artificial Intelligence JavaScript Postman