Data Engineer (Databricks)
Avensys Consulting · Singapore
Avensys is a reputed global IT professional services company headquartered in Singapore. Our service spectrum includes enterprise solution consulting, business intelligence, business process automation and managed services. Given our decade of success we have evolved to become one of the top trusted providers in Singapore and service a client base across banking and financial services, insurance, information technology, healthcare, retail, and supply chain.We are currently looking to hire Data Engineer (Databricks).This is an exciting opportunity to expand your skill set, achieve job satisfaction and work-life balance. More details as below.Job Role:Job DescriptionThe Data Engineer will be responsible for designing, developing, and maintaining scalable and reliable data pipelines on Databricks and cloud platforms. The role requires integrating diverse data sources, ensuring high-quality data processing, and supporting analytics, reporting, and machine learning workloads.The role involves collaborating closely with analytics, product, and infrastructure teams to enhance the company’s data platform while adhering to best practices for governance, monitoring, and reliability. What will you do?- Develop and maintain ETL pipelines for centralized data storage systems (e.g. Delta Lake).- Integrate data from databases, APIs, log files, streaming platforms, and external providers- Develop data transformation routines to clean, normalize, and aggregate data- Apply data processing techniques to handle complex or inconsistent datasets- Contribute to frameworks and best practices for code development and deployment- Implement data governance in alignment with company standards- Partner with analytics and product leaders to design and operationalize pipelines- Collaborate with infrastructure leaders to advance cloud-based data platforms- Explore new tools and techniques leveraging Azure, Databricks, or related platforms- Monitor data pipelines to detect and resolve issues promptly- Develop monitoring tools, alerts, and automated error-handling mechanisms- Analyse business requirements and identify data extraction requirements- Attend requirement grooming and refinement sessions with users- Develop and maintain ETL pipelines for ingestion, transformation, validation, and loading- Optimise performance and batch scheduling- Develop dashboards, reports, scorecards, and data visualizations- Perform SIT, data validation, data profiling and confirm data accuracy- Validate completeness and consistency of ETL Loads- Support UAT and production implementationQualifications The ideal candidate should possess:- 3 or more years of experience in data engineering with scalable pipelines- Strong experience designing data solutions including data modelling and distributed computing architectures- Hands-on experience with data processing jobs using PySpark, Spark SQL, and Databricks notebooks/jobs- Experience orchestrating data pipelines with ADF, Airflow, or similar tools- Experience with both real-time and batch data processing- Experience building pipelines on Azure, with AWS experience beneficial - Proficiency in SQL including window functions and performance optimization- Understanding of DevOps tools, Git workflows, and CI/CD pipelines- Familiarity with Scrum methodology and experience working in Scrum teams- Ability to apply Scrum practices in a practical project context- Strong problem-solving and collaborative mindset- Experience with streaming technologies such as Apache Kafka, Apache Flink, or AWS Kinesis- Ability to design and implement real-time data processing pipelines Must have Databricks Certified Data Engineer Associate and Databricks Certified Data Engineer Professional is strongly preferredCONSULTANT DETAILSConsultant Name: Abinaya R Reg No: R1765546 Avensys Consulting Pte Ltd EA Licence 12C5759 Privacy Statement: Data collected will be used for recruitment purposes only. Personal data provided will be used strictly in accordance with the relevant data protection law and Avensys' privacy policy.