Highlights:
5.00 – 10.00 Years
14.00 – 20.00 INR (Lacs)/Yearly
Full-time
Gurgaon, Noida, Hyderabad
Skills
Data Engineer
Data Engineering
Python
AI
Pyspark
AWS
GIT
Bash
Bash/Shell scripting
Roles & Responsibility
We have an opening of one Data Engineer profile to fill in. JD is given below, please share compatible profiles.
4+ years of overall IT experience, including hands-on experience with Big Data technologies.
Strong hands-on experience in Python and PySpark.
o Python should be used extensively for application development, ETL (Extract, Transform, Load) processes, and Data Lake curation.
Experience in building PySpark applications using Spark DataFrames in Python with development tools such as Jupyter Notebook and PyCharm (IDE).
Proven experience in optimizing Spark jobs that process large-scale datasets and high-volume workloads.
Hands-on experience with version control systems, particularly Git.
Experience working with AWS Analytics services, including:
o Amazon EMR
o Amazon Athena
o AWS Glue
Experience with AWS Compute and Storage services, including:
o AWS Lambda
o Amazon EC2
o Amazon S3
o Other AWS services such as Amazon SNS
Working knowledge of Bash/Shell scripting is an added advantage.
Requirements
Experience in designing and building ETL pipelines to ingest, copy, cleanse, and transform data across multiple formats, including:
o CSV
o TSV
o XML
o JSON
Experience working with various file formats, including:
o Fixed-width files
o Delimited files
o Multi-record file formats
Good understanding of Data Warehouse concepts, including:
o Facts and Dimensions
o Star Schema
o Snowflake Schema
Experience with columnar storage formats, such as:
o Parquet
o Avro
o ORC
Strong knowledge of data compression techniques, including:
o Snappy
o Gzip
Good to have experience with at least one AWS database service:
o Amazon Aurora
o Amazon RDS
o Amazon Redshift
o Amazon ElastiCache
o Amazon DynamoDB
Good to have experience with AI/ML.
