Source description
About the role
We are seeking a motivated and data-savvy Pyspark Bigdata Developer r to join our growing team. This role is perfect for an individual passionate about data, with foundational skills in both software development and quality assurance. The successful candidate will be responsible for developing and testing data-centric applications, with a strong focus on writing complex Strong programming skills in PySpark and Python in a BigData environment. This hybrid role offers a unique opportunity to work across the entire data lifecycle, from development and testing of data pipelines to delivering actionable insights through BI tools.
Responsibilities
Strong programming skills in PySpark in a BigData environment. Familiarity with big data processing tools and techniques. Experience with the Hadoop ecosystem including Hive, HDFS, Sqoop, Spark, Impala, Scala, etc.,Well versed with shell scripting and Autosys scheduler. Good understanding of distributed systems. Should be familiar with data warehouse concepts. Experience with streaming data platforms. Excellent analytical and problem-solving skills. Experience with writing complex SQL queries. Knowledge on data modeling and data design is essential. Should be independent and resourceful dealing with risks/issues and resolving them in a timely manner. Excellent communication and articulation skills. Skills Pyspark : Strong programming skills in PySpark in a BigData environment. SQL: Strong proficiency in writing and optimizing complex SQL queries. Experience with a major relational database system (e.g., SQL Server, Hive, Impala ETL Concepts: Basic understanding of ETL (Extract, Transform, Load) processes and data pipeline concepts. Testing: Familiarity with data quality and testing methodologies. Version Control: Experience with version control systems, such as Git.
Qualifications
Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent practical experience. Minimum 4 years of experience in a role involving data development, database management, or software quality assurance. Strong understanding of relational databases and data warehousing concepts. Foundational knowledge of at least one programming language (SQL,Pyspark,Python)
Education: Bachelor’s degree/University degree or equivalent experience
This job description provides a high-level review of the types of work performed. Other job-related duties may be assigned as required.
Job Family Group: Technology ------------------------------------------------------ Job Family: Applications Development ------------------------------------------------------ Time Type: Full time ------------------------------------------------------ Most Relevant Skills Please see the requirements listed above. ------------------------------------------------------ Other Relevant Skills For complementary skills, please see above and/or contact the recruiter. ------------------------------------------------------ Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.
If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi .
View Citi’s EEO Policy Statement and the Know Your Rights poster.
More at Citi
Related open roles
Java Spark Developer – Assistant Vice President
India · Hybrid
Senior Algo Trading Software Engineer (VP)
United Kingdom · Hybrid
Senior Software Engineer (Assistant Vice President) – Margin Technology
United States
Java Lead Vice president
India
Senior Java Lead Software Engineer
Canada · Hybrid
Officer - Data Engineer (Big Data, Python, PySpark)
India