AI Data Platform Lead

Newbridge Alliance · Singapore

Sector
AI
Function
Product & Engineering
Level
Lead
Employment type
Full Time
Posted
2026-08-13
Source
mycareersfuture

Our client is a high-scale global consumer platform serving millions of users daily. They are rebuilding their data foundation as an AI-native platform to power the next generation of programmatic advertising, personalized discovery, and marketplace intelligence.We are hiring a hands-on Staff / Lead Data Engineer (AI-Native) to architect petabyte-scale systems that feed real-time bidding, recommendation models, and LLM-powered analytics products.Role Mandate: 80% Hands-on Engineering / 20% Technical LeadershipThis is NOT a pure management role. You will be a player-coach - designing and shipping critical pipelines yourself while mentoring a small pod of engineers. Strong individual contributors with no formal people management experience but with deep technical depth are strongly encouraged to apply.What You Will Build1. AI-Native Data Foundation at ScaleDesign, build, and operate AI-ready batch + streaming ETL/ELT pipelines ingesting 100TB+ daily from ad servers, mobile SDKs, transactional systems, and 3rd-party APIs. Build for LLM and ML consumption from day one.2. Real-Time & Agentic Data SystemsDevelop low-latency streaming jobs using Spark Streaming, Flink, or Kafka Streams for real-time use cases: fraud detection, bid optimization, dynamic pricing, and real-time personalization. Enable online inference and agentic decisioning.3. Lakehouse for AI & AnalyticsModel and optimize massive datasets on a modern lakehouse [Databricks / Snowflake / BigQuery] to serve BI, embedded analytics, and AI/ML workloads with a focus on performance, cost, and feature freshness.4. Data Products for AIBuild reliable data products for Data Science & ML: feature stores [Feast / Tecton], vector stores for semantic search & RAG, training datasets with point-in-time correctness, and online-offline parity.5. Reliability for Tier-0 AI SystemsOwn data quality, observability, anomaly detection, and lineage for Tier-0 datasets that directly power revenue, user experience, and model performance.Required Qualifications8+ years building large-scale distributed data systems for programmatic advertising, digital media, marketplaces, or large consumer internet platforms, supporting 50+ downstream engineers / analysts / scientistsExpert-level SQL and strong production coding in Python, Scala, or JavaDeep hands-on experience with distributed processing: Spark, Kafka, Flink or equivalentProven expertise with cloud data platforms: Databricks, Snowflake, BigQuery, Redshift, AWS / GCP / AzureStrong data modeling for AI: dimensional, Data Vault, lakehouse, medallion architecture optimized for analytics and MLTrack record handling high-volume, semi-structured, late-arriving event data at TB-PB scalePreferred - You Stand Out If You Have:AI-Native Data Engineering: Experience building AI-native infrastructure - feature platforms, vector DBs, LLM eval pipelines, RAG data pipelines, unstructured data processing for GenAI, AI agent memory/context infrastructure.AdTech Domain: DSP/SSP internals, impression/click/conversion pipelines, SKAN, MMM/MTA, Conversions API, identity resolution & signal loss mitigationEcommerce / Marketplace Domain: Product catalog & taxonomy at scale, pricing experimentation, inventory forecasting, seller analytics, search & recommendation dataML Data Infra: Feature Stores, online-offline parity, training data infrastructure, model monitoring data loopsLeadership & Impact: Experience leading 0-to-1 architecture for critical domains like real-time advertising or attribution. Demonstrated impact on cost optimization of $1M+ in annual cloud spend.Privacy & Trust: Knowledge of privacy-enhancing tech: differential privacy, data clean rooms, secure multi-party compute.Interested candidates, email latest resume to [email protected]

Apply on mycareersfuture →
AI Distributed Processing Microsoft Azure Digital Media Scala AWS Advertising Platform SQL Programming