Data Engineering Intern
A 6-month, in-office internship building Palladium — the data + AI platform trusted by hundreds of mid-to-large Indian enterprises. You'll work directly with the founders and engineering team to acquire, clean, and process large-scale data that powers intelligent search and decision-making. Core work: web scraping and API reverse-engineering, Python data pipelines, PostgreSQL, and preparing data for AI/ML workflows.
Responsibilities
- Build crawlers and scrapers (e.g. with Selenium) and reverse-engineer web APIs to acquire data at scale.
- Design and build Python data pipelines to extract, clean, transform, and process large, messy datasets.
- Model and query data in PostgreSQL — schema design, optimization, and handling large tables.
- Prepare and structure datasets that feed machine-learning models and AI-driven analytics.
- Instrument data quality and reliability across ingestion and processing.
- Deploy and run pipelines on Linux servers, and support monitoring and uptime.
- Work closely with founders and engineers to turn raw data into product features used by Indian enterprises.
Requirements
- Currently pursuing or recently completed a B.Tech in Computer Science or a related field.
- Strong programming skills in Python.
- Comfort working with data — cleaning, transforming, and analysing large datasets (e.g. pandas, SQL).
- Basic understanding of PostgreSQL or relational databases.
- Familiarity with web scraping, APIs, or data-processing tools is a strong plus.
- Exposure to Linux environments.
- Genuine interest in data engineering, pipelines, and AI/ML data workflows.
- Strong analytical and problem-solving skills with attention to detail.
Perks
- Hands-on experience building production data pipelines and infrastructure.
- Work with large-scale, real-world datasets from across India's economy.
- Exposure to AI/ML data workflows and intelligent data products.
- Work directly with the founders in a fast-growing SaaS startup.
- Build technology used by leading Indian enterprises.