AI Engineer – LLM Algorithm Engineer (Agentic Commerce)
Shopee Ip Singapore · Singapore
Job Description:Lead continuous pre-training and post-training of large language models (Dense and MoE architectures) for vertical business domains, including domain-specific data synthesis, multi-stage data curation, and monthly iteration to support downstream business integration.Design and build autonomous agents (e.g., Plan-and-Execute agents) that automatically retrieve, extract, and synthesize domain knowledge from large-scale multilingual/unstructured data sources, producing high-confidence corpora for continued pre-training.Develop and iterate RL-based post-training methods (e.g., CPT, SFT, DAPO) to improve model performance on domain-specific QA, reasoning, and knowledge-modeling tasks.Build and optimize multimodal LLM applications for content generation and understanding, including product copywriting, semantic search recall, and business potential/CTR prediction.Design end-to-end agent workflows integrating query understanding, content generation, validation, and downstream product/search systems.Build generative representation-learning pipelines (e.g., token-based generative pre-training on behavioral sequences) to produce reusable embeddings for downstream ranking and recommendation models.Collaborate cross-functionally to translate business requirements into scalable model training and agent system design; continuously monitor production metrics to guide model and agent optimization.Requirements:Master's degree or above in Computer Science, Natural Language Processing, Artificial Intelligence, Electrical and Electronics Engineering, Signal Processing or a related field. Minimum 5 years of full-time industry experience in LLM/ML algorithm engineering, including hands-on experience with large-scale model continuous pre-training (Dense and/or MoE architecture) for vertical business domains.Hands-on experience designing and building autonomous agent systems (e.g., using LangGraph or similar frameworks) for automated knowledge retrieval, extraction, and synthesis pipelines, combined with hands-on experience in LLM post-training techniques (SFT, RL-based methods such as DAPO) and data-mixture optimization techniques for pre-training data curation.Experience fine-tuning and deploying multimodal large models for content generation, semantic recall, or business potential prediction, with demonstrated production impact on key business metrics (e.g., CTR, conversion).Good programming and engineering skills; solid foundation in classic ML techniques (e.g., LightGBM, XGBoost, Bayesian modeling) is a plus.Good problem-solving skills; able to independently drive projects from research through to large-scale production deployment.Prior experience in multimodal retrieval/recommendation systems (e.g., cross-modal contrastive learning for content matching) will be a strong plus.