Data Engineering Intern
Apollo is building out a data lake using an already-established medallion architecture (Bronze / Silver / Gold), and expanding it as we bring on new clients and migrate existing ones. A large part of this work is moving business logic out of Tableau workbooks and into a centralized, documented Gold layer.
We're looking for a Data Engineering Intern to support that effort. You'll work alongside our data engineering team on pipeline development, SQL implementation, and — importantly — the discovery work that has to happen before anything gets migrated: figuring out how a report actually calculates a number today, writing an equivalent SQL definition, and documenting it so it can be validated and standardized.
This is a hands-on role. You won't be setting technical direction, but you will be doing real work that ships, with your output reviewed by senior engineers before it reaches clients.
Overview
Location: On-site/remote (White Plains, NY)
Type: Part-time (20 hours per week)
Compensation: $2500 monthly
Responsibilities
- Reverse-engineer existing Tableau reports to identify how metrics are calculated — calculated fields, LOD expressions, filters, and context filters
- Contribute to on external data imports as part of client management & onboarding
- Write SQL implementations of those metrics and compare outputs against Tableau to confirm they match
- Document metric definitions using our standard templates so report owners can confirm or correct proposed logic
- Write and optimize complex SQL queries against large datasets; Write and test SQL queries against Gold layer views and tables
- Build web scraping tools to source external data as needed
- Support ETL pipeline development in Python under the direction of the data engineering team; Help maintain extraction and metadata tooling used across client environments
Requirements
- Currently pursuing a degree in Computer Science, Data Science, Information Systems, Statistics, or a related field
- Working SQL knowledge, including joins, aggregations, and window functions
- Python experience, including familiarity with pandas
- Comfort reading and reasoning about code or data structures you didn't write
- Attention to detail — much of this work is confirming that two implementations produce identical results
- Ability to communicate clearly in writing, including with non-technical stakeholders
- Able to work 20 hours per week
Nice to Have
- Exposure to Tableau or Power BI, whether as a builder or a user
- Familiarity with cloud data platforms, ideally Azure Synapse or Azure Data Lake
- Experience with Git / Github
- Coursework or projects involving ETL or data pipelines
- Microsoft 365 experience
Our Evaluation Method
Our process includes a practical component — we're interested in how quickly you can get up to speed on an existing database/data lake architecture, and we'll ask you to work through some complex SQL queries.
Please submit your resume and a cover letter to info@apollov2.com.
ApolloV2 is an equal opportunity employer with a commitment to hiring people with diverse backgrounds. We do not discriminate based on age, civil or family status, disability, ethnicity, gender, race, religion, or sexual orientation.
