Data Engineering & Analytics Engineer (Public Sector)
Websparks · Singapore
1-year contract, renewableGovernment projectHybrid work arrangementWe are looking for a Data Engineering & Analytics Engineer to design, build, and operate the data pipelines and data products supporting the future SSOE platform.You will transform data from enterprise systems, operational platforms, applications, and infrastructure into trusted and usable data products for applications, operational reporting, analytics, capacity planning, and decision-making.The role spans on-premise, GCC and hybrid environments, with an emphasis on production-grade data engineering using cloud-native and managed data capabilities.What You Will Be Working OnAs a Data Engineering & Analytics Engineer, you will own the data lifecycle from source systems through ingestion, transformation, modelling, quality, and serving.You will build pipelines that extract and ingest data from enterprise and operational systems, transform it into consistent and trusted datasets, and make that data available to applications, dashboards, reporting, analytics, and machine-learning use cases.You will work closely with the Logging & Data Platform Engineer on shared platform capabilities and with Software Engineers and other consumers to define reliable data interfaces and products.Key ResponsibilitiesData Pipeline EngineeringDesign, build, and operate production-grade data pipelines for data extraction, ingestion, transformation, and loading (ETL/ELT)Integrate data from on-premises systems, enterprise applications, APIs, databases, SaaS platforms, files, streams, cloud services, and other operational data sourcesDevelop batch, incremental, change-data-capture (CDC), streaming, and event-driven ingestion patterns based on source-system and business requirementsBuild transformation pipelines that clean, enrich, standardise, join, aggregate, and structure raw data into trusted datasetsDesign secure and resilient mechanisms for transferring and synchronising data between on-premises, GCC, AWS, Azure, and other approved environmentsDesign pipelines for failure handling, retry, recovery, idempotency, scalability, and changing data volumesAutomate pipeline deployment, configuration, testing, and operationData Architecture & ModellingDesign and maintain cloud-native and hybrid data stores, data lakes, and analytical datasetsDevelop data models that provide consistent representations of enterprise, operational, and asset informationDefine schemas and data contracts between data producers and downstream consumersDesign data structures appropriate for operational applications, reporting, analytics, and machine-learning workloadsApply backwards-compatible schema changes and coordinate changes that may affect downstream consumersMaintain data lineage and metadata so datasets are traceable and discoverableWork with platform and application teams to define appropriate data-serving and integration patternsData Quality & ReliabilityImplement automated data validation, reconciliation, completeness, consistency, and quality controls throughout the pipeline lifecycleMonitor data freshness, pipeline health, processing latency, and data-quality indicatorsDetect and investigate ingestion failures, source-system changes, data-quality anomalies, and reconciliation differencesPrevent invalid or incomplete data from silently propagating to downstream consumersDefine appropriate SLOs for data freshness, availability, and pipeline reliabilityBuild monitoring, alerting, error handling, and recovery into data pipelines from the outsetAnalytics & Data ProductsBuild trusted datasets and reusable data products for applications, dashboards, operational reporting, and analyticsDevelop datasets supporting asset intelligence, operational visibility, capacity planning, trend analysis, and decision-makingEnable advanced analytics and machine-learning use cases using cloud-native data, analytics, and AI/ML capabilitiesWork with users and stakeholders to translate operational questions into appropriate datasets, metrics, and analytical productsSupport exploratory analysis and prototyping where required before operationalising successful approachesEnsure analytical outputs are based on governed, traceable, and reproducible dataData IntegrationDesign data architectures spanning on-premise infrastructure and cloud platformsIntegrate traditional enterprise systems with modern cloud-native data capabilitiesDesign for connectivity constraints, network boundaries, security zones, and data-residency requirementsImplement appropriate buffering, checkpointing, retry, and reconciliation where data crosses environment boundariesSelect appropriate integration patterns based on data volume, latency, source-system capability, and operational requirementsWork with infrastructure, network, security, and platform teams to establish secure data flowsSecurity & GovernanceEnsure data is collected, transmitted, stored, processed, and accessed according to applicable security requirementsEnforce appropriate access controls and least-privilege principles for data platforms and pipelinesEnsure sensitive information is appropriately classified and protected throughout the data lifecycleMaintain auditability and traceability of data-processing activitiesApply retention, archival, lifecycle, and deletion requirements to data productsParticipate in security, architecture, data-governance, and operational-readiness reviewsReliability & OperationsOperate and support production data pipelines and data productsParticipate in operational support and on-call responsibilities for owned servicesInvestigate production incidents and contribute to root-cause analysis and preventative improvementsMonitor pipeline performance, capacity, reliability, and costMaintain architecture documentation, data definitions, operational procedures, and runbooksContinuously improve pipeline automation, reliability, performance, and maintainabilityWhat We Are Looking ForExperienceMinimum 3–5 years of experience in data engineering, cloud data engineering, analytics engineering, software engineering, or a related disciplineAt least 2 years of hands-on experience designing, building, and operating production-grade data pipelinesDemonstrated experience with data extraction, ingestion, ETL/ELT, transformation, data modelling, and data qualityExperience using AWS and/or Azure native data capabilitiesExperience integrating data from APIs, databases, enterprise systems, files, or streaming sourcesExperience implementing batch, incremental, CDC, and/or event-driven data pipelinesExperience working with on-premises and/or cloud environments, with an understanding of hybrid integration patternsExperience applying software-engineering practices such as version control, automated testing, CI/CD, monitoring, and Infrastructure as Code to data solutionsTechnical SkillsCloud: AWS/Azure-native logging, streaming, storage, search and data servicesOn-Premise: Enterprise servers, networks, applications, databases, virtualised infrastructure, and log sourcesData Engineering: Python, SQL, ETL/ELT, batch, incremental, CDC, streaming, and event-driven patternsData Modelling: Relational, dimensional, analytical, and domain-oriented data modellingData Quality: Validation, reconciliation, quality monitoring, lineage, and anomaly detectionInfrastructure as Code: Terraform / OpenTofuCI/CD: GitLab CI/CD, SHIP-HATS or equivalent automated deployment practicesAnalytics: Data preparation, analytical datasets, reporting, statistical analysis, and ML enablementEngineering PracticesTreats data pipelines and data products as production software, not one-off scriptsKeeps pipeline code, schemas, infrastructure, and configuration under version controlUses automated testing, CI/CD, Infrastructure as Code, and monitoringDesigns pipelines for failure, retry, idempotency, scalability, and changing workloadsValidates data at ingestion and transformation boundariesEstablishes explicit data contracts between producers and consumersUnderstands when to use managed cloud-native capabilities rather than unnecessarily building and operating infrastructureConsiders downstream consumers before making schema or behavioural changesAutomates repeatable data-processing and operational activitiesBalances technical excellence with pragmatic delivery and operational sustainabilityNice to HaveExperience with Singapore Government platforms such as TechPass, SHIP-HATS, SEED, and GCCFamiliarity with OC/SN data-classification requirementsAWS or Azure cloud certificationsExperience designing data architectures spanning on-premises and cloud environmentsExperience building data products consumed by applications, dashboards, operational teams, or leadershipExperience with managed cloud analytics and AI/ML capabilitiesExperience with data lineage, metadata management, and data cataloguingExperience working with enterprise asset management, MDM, network, procurement, or operational systems