Data Engineer
Data Engineer designs, builds, and manages secure, scalable data processing systems and pipelines to turn raw data into actionable insights.
About the Role
Programming & Querying: Strong proficiency in SQL for data analysis and relational databases, alongside Python or PySpark for automation and pipeline scripting. Core Tools: Deep understanding of BigQuery (data warehousing/analytics), Cloud Storage (GCS) (blob storage), and Pub/Sub (real-time messaging). Processing & Orchestration: Expertise in Dataflow (batch/streaming via Apache Beam), Dataproc (managed Spark/Hadoop), and Cloud Composer (workflow orchestration via Apache Airflow). Modern Data Stack: Familiarity with transformation tools like dbt, Dataform, and data governance frameworks like Dataplex.
Responsibilities
- →Key Responsibilities Logic Discovery And Requirements Analysis
- →Work with pricing analysts to identify custom definitions, calculations, transformations, dependencies, assumptions, edge cases, and ambiguities currently implemented in the Gold layer.
- →Reverse-engineer analyst-written SQL, spreadsheet logic, and ad hoc calculations to define intended business behaviour.
- →Translate legacy logic into clear technical requirements and migration plans.
- →Migration to the Silver Layer
- →Re-implement approved pricing definitions as governed, reusable, documented dbt models in the Silver layer.
- →Follow data modelling, naming, testing, version control, and code review standards.
- →Build maintainable, scalable, auditable models independent of undocumented analyst knowledge.
- →Apply dbt tests for uniqueness, not-null, accepted values, relationships, and business rules.
- →Parity Validation and Quality Assurance
- →Compare migrated Silver-layer results with existing Gold-layer outputs and resolve differences.
- →Validate results across representative periods, segments, boundary conditions, and known exceptions.
- →Confirm migrations only after parity review and pricing analyst approval.
- →Analyst Collaboration and Sign-Off
- →Partner with pricing analysts to clarify calculations, confirm intended behaviour, and align on expected results.
- →Lead walkthroughs and reviews; document analyst approval before retiring legacy definitions.
- →Communicate risks, open questions, dependencies, and decisions to technical and business stakeholders.
- →Documentation and Lineage
- →Document business meaning, calculation rules, assumptions, exceptions, source inputs, ownership, lineage, tests, and downstream consumers in dbt and Confluence.
- →Ensure documentation enables future engineers and analysts to maintain logic without relying on the original analyst.
- →Gold-Layer Cleanup and Cutover
- →Decommission ad hoc Gold-layer logic after the Silver replacement is validated, approved, adopted, and downstream dependencies are updated.
- →Verify retired logic is no longer used and operational documentation reflects the new source of truth.
Requirements
- →5–8 years of data engineering experience on analytical data platforms.
- →Strong SQL skills: joins, window functions, CTEs, aggregations, conditional logic, and optimisation.
- →Hands-on dbt experience across models, tests, documentation, and lineage.
- →Strong Snowflake experience, including modelling, performance tuning, and warehouse concepts.
- →Experience refactoring, migrating, or modernising complex analytical logic.
- →Ability to untangle ad hoc SQL and spreadsheet calculations and validate outputs.
- →Strong understanding of data quality, reconciliation, testing, and release controls.
- →Skilled at working with non-engineering stakeholders to clarify requirements and validate outcomes.
- →Excellent communication and documentation habits.
