AI Engineer
Replicast AI · Remote / APAC
About the role We’re building an AI Digital Human SaaS platform that enables lifelike avatars to listen, understand, respond using enterprise knowledge, and communicate naturally in real time. The platform is designed to run entirely on our own infrastructure, ensuring data remains secure within our network. You’ll take end-to-end ownership of the core system — from knowledge and AI models to speech, real-time interactions, and the digital human experience.
What you’ll do Build and run the real-time avatar (rendering, lip-sync, expressions) in the browser. Self-host language models and build retrieval over our private knowledge base. Set up on-prem speech: speech-to-text and text-to-speech. Keep responses fast and conversations natural (low latency, interruptions). Deploy and run the models on our GPUs, fully self-contained.
Required skills Strong full-stack engineering (application and backend) Real-time browser 3D/graphics and real-time media streaming Self-hosting large language models on local infrastructure (no cloud APIs) Retrieval-augmented generation (RAG) over private data with a self-hosted vector database On-prem speech: speech-to-text and text-to-speech Real-time avatar animation: lip-sync and facial expressions Low-latency streaming and conversation handling GPU infrastructure: containerization, model serving, on-prem deployment Hardware sizing: VRAM and GPU budgeting across LLM / STT / TTS, and spec’ing the machine to match Comfortable across application and ML/model-serving languages
Nice to have Photoreal / neural avatars Model optimization (quantization, batching) for limited GPUs On-prem / air-gapped security and compliance