AI Infrastructure Engineer

Kaishi Partners · Singapore

Sector
AI
Function
Product & Engineering
Level
Mid-Level
Employment type
Full Time
Posted
2026-09-20
Source
mycareersfuture

Our client is an early-stage, venture-backed AI startup building a personal agent that understands users’ preferences, helps them discover relevant products, and follows through on their decisions. Combining AI agents, personalisation, and commerce, the team aims to create a service that learns from real customer experiences and becomes more useful over time.They are expanding their engineering team in Singapore, offering early hires substantial ownership of the product and technical architecture, working closely with the founders across AI infrastructure and full-stack development. You will be the first Infrastructure Engineers to join their teamThe opportunityYou will build the infrastructure that makes AI agents dependable and affordable to run. These agents need to retain context across conversations, execute actions reliably, recover from interruptions, and connect their decisions to outcomes that may arrive weeks later.This is a hands-on software engineering role with substantial ownership of architecture and production reliability. You will work closely with product engineers and researchers to turn emerging agent capabilities into systems people can rely on.What you’ll doBuild and evolve the Python infrastructure supporting agent execution, persistent state, model routing, and tool use.Design recovery mechanisms for interrupted workflows, including retries, failure handling, and protection against duplicate actions.Develop event and execution-history pipelines that connect conversations, agent decisions, actions, and outcomes.Establish deployment, monitoring, and debugging practices that help a small team operate effectively.Improve latency, reliability, and model usage costs as the product develops.Make pragmatic architectural decisions and take ownership of systems through production failures and changing requirements.What you’ll bringStrong software engineering skills in Python and experience building distributed backend systems.Experience owning production infrastructure, investigating incidents, and making reliability trade-offs.Practical understanding of asynchronous processing, state management, APIs, and data persistence.The ability to balance immediate delivery with sound technical foundations.Comfort working independently in a small team where requirements and architecture are still evolving.Useful additional experienceAgent runtimes, model API integrations, or ML infrastructure.Durable workflows, event-driven systems, or execution tracing.Measuring and optimising inference cost and latency.To applyShare your profile and a system you helped build, including the architectural decisions you owned and what you learned from operating it.

Apply on mycareersfuture →
AI Product Costing and Pricing Reliability Analysis Agents Pipeline Development Team Facilitation Monitoring Workflow Design