We are seeking a highly skilled Data Engineer to join our dynamic team. The ideal candidate will have strong experience in building data pipelines, working with modern data platforms, and implementing scalable cloud-based solutions. This role requires hands-on expertise in SQL, Python, PySpark, and AWS, along with solid understanding of data modelling and transformation frameworks
Key Responsibilities
- Design, develop, and optimize scalable data pipelines and ETL workflows.
- Build and manage data models using DBT and integrate datasets for analytics and reporting.
- Work with Snowflake to write efficient SQL queries and manage data warehousing activities.
- Develop distributed data processing solutions using PySpark on Databricks.
- Implement data ingestion frameworks using Python, APIs, and AWS services.
- Collaborate with cross-functional teams to understand data requirements and deliver solutions.
- Ensure data quality, reliability, and performance across pipelines.
- Deploy and maintain data workflows using CI/CD and version control best practices.
Mandatory Skills
- Strong expertise in SQL (Snowflake) — minimum 2 years
- Hands-on experience with DBT — minimum 2 years
- PySpark (Databricks) — minimum 2 years
- Python with solid understanding of OOP concepts — minimum 2 years
- AWS (S3, ECS, SQS, Lambda, EMR) — minimum 2 years
- Working knowledge of CSV, JSON, Parquet file formats and API integrations using Python
- Good communication skills
Secondary / Good-to-Have Skills
- Experience with Sigma or Tableau for data visualization
- Linux CLI proficiency
- Python web scraping experience
- Exposure to Datadog for monitoring
- Experience with GitHub for version control and CI/CD
- Knowledge of Terraform
- Computer Science degree from a reputed academic institution