Data Engineer
Skills
About the role
This is a hands-on data engineering role supporting large-scale data pipelines and infrastructure for a company that crawls and delivers web data to AI training teams. It suits an engineer comfortable with distributed systems, web scraping at scale and production data workloads. The role is fully remote worldwide but requires schedule overlap with EST business hours.
What you’ll do
- Maintain, optimize and troubleshoot database queries and related data systems
- Build and improve data pipelines that collect, process, transform, validate and deliver large-scale datasets
- Support web scraping and data collection initiatives including scripts and tooling
- Monitor and troubleshoot pipeline issues and data quality concerns
- Document engineering work including queries, pipelines and technical decisions
- Participate in R&D projects to improve data products and workflows
What they’re looking for
- Bachelor's degree or equivalent work experience
- Advanced Python with async programming and multiprocessing experience
- Hands-on high-volume web scraping experience, including proxies, rate limiting and anti-bot evasion
- Experience designing distributed pipelines using task queues such as Celery, Kafka or RabbitMQ
- Practical experience with columnar/analytical warehouses, ideally Databend, ClickHouse or BigQuery
- Docker and Kubernetes experience, including Helm charts and autoscaling
- Linux and bare-metal operations experience
- CI/CD experience with tools such as GitHub Actions or ArgoCD
Nice to have
- Experience with platform APIs and large media/metadata datasets
What’s on offer
- Competitive salary, benefits and equity package
- Fully remote team
Questions about this role
Is this role remote?
Yes, it is fully remote worldwide, but requires overlap with EST business hours.
Do I need a degree?
A bachelor's degree or equivalent work experience is required.
How this role compares
One of 11 open Data Engineer roles we are tracking today.
Posted 7 days ago, with 7 of the others posted since and 3 still open from before.
Related roles
Lead Data Engineer – Databricks & AI | Remote Canada
Lead data engineer role for a financial services client, requiring 10+ years of Databricks expertise and experience collaborating with AI/ML teams to build scalable data infrastructure.
Data Analyst - Data Quality & Databricks
A data analyst role centered on data quality, profiling and reconciliation using Databricks, SQL and PySpark, open across Bangalore, Chennai and Pune for candidates with 7-11 years of experience.
Data Engineer, Data Platform & ML (Remote)
A Poland-remote Data Engineer role building a cloud-based corporate data warehouse for business analytics and machine learning. It suits a data engineer with at least three years of experience in Python, SQL, warehouse design, and modern data tooling.
Data Engineer, Data Platform & Machine Learning (Remote)
A Serbia-remote data engineering role building Tabby's corporate data warehouse, data integrations, and pipelines. It is intended for a data engineer with warehouse design, Python, SQL, and cloud data-stack experience.
Senior Data Engineer, Data Platform & ML (Remote)
A remote senior data engineering role in Yerevan focused on feature-store development, data services, streaming, and production data pipelines. It is suited to an engineer with scalable systems experience across data, machine learning, or backend development.
Data Architect / Data Engineering Lead
A principal data architecture and engineering leadership role at UnitedHealth Group's Optum/LHI, based in Minnetonka, Minnesota. Suited to a senior data leader with deep Snowflake/Databricks and Airflow experience who wants to combine architecture ownership with hands-on delivery.