Data Engineer
Apollo is building out a data lake using an already-established medallion architecture, and expanding it as we bring on new clients and migrate existing ones. We're looking for a Data Engineer to lead this effort — owning the data lake's architecture and health, leading external data imports as we onboard new clients, and leading all API-based data imports, which flow directly into the data lake rather than our transactional database.
This is a leadership role on the data side: you'll be the person setting direction on how data flows in, how it's structured, and how ETL pipelines are built and maintained, while also staying hands-on with the pipelines and queries themselves.
Overview
Location: On-site (White Plains, NY)
Type: Full-time
Compensation: $80,000 - $90,000
Responsibilities
- Lead the data lake build-out and expansion using our existing medallion architecture, onboarding both new and existing clients
- Own external data imports as part of client onboarding
- Lead all API-based data imports, ensuring API data is imported directly into the data lake rather than the transactional database
- Design, build, and maintain ETL pipelines using Python
- Write and optimize complex SQL queries against large datasets
- Build web scraping tools to source external data as needed
- Set technical direction and best practices for the team's data engineering work
Requirements
- Experience with Azure Synapse (data lake)
- Experience with Azure Data Lake Studio
- Professional SQL experience, preferablyin a leadership role
- Python experience, preferrably in production
- Experience building and maintaining ETL pipelines
- ETL experience using Python
- Web scraping experience
- Comfort owning deployments, not just writing code
- Experience with Git/GitHub, Docker, or model deployment tools
- Ability to clearly communicate insights to both technical and non-technical stakeholders
- Ability to work independently and collaboratively in a fast-paced environment
- Able to work 40 hours per week, in-office
Nice to Have
- Experience with PySpark
- Experience with Kafka
- Experience with Flink
- Microsoft 365 experience
Our Evaluation Method
Our process includes a practical component — we're interested in how quickly you can get up to speed on an existing database/data lake architecture, and we'll ask you to work through some complex SQL queries.
Please submit your resume and a cover letter to info@apollov2.com.
ApolloV2 is an equal opportunity employer with a commitment to hiring people with diverse backgrounds. We do not discriminate based on age, civil or family status, disability, ethnicity, gender, race, religion, or sexual orientation.
