Staff Data Engineer
Skills
About the role
This is a remote, full-time Staff Data Engineer position for a U.S.-based data engineer with at least seven years of experience building streaming and event-driven systems. You’ll establish Harbor Compliance’s end-to-end data platform, connecting operational, financial, CRM, and HR data into an AI-ready source for decision-making. The role is a foundational hands-on hire focused on real-time pipelines, vector infrastructure, reliability, and data products.
What you’ll do
- Design and own near-real-time pipelines using streaming, CDC, and event-driven patterns.
- Build vector-database and embedding infrastructure for semantic retrieval and AI-agent use cases.
- Ingest platform, financial, CRM, and HRIS data for batch and real-time consumers.
- Help architect the warehouse or lakehouse supporting the data platform.
- Create transformation layers for analytics-ready datasets.
- Implement monitoring, failure alerts, and lineage tracking.
- Develop foundations for self-service and AI-supported reporting.
- Establish governance documentation and access controls.
What they’re looking for
- Have seven or more years of hands-on data engineering experience.
- Have built production streaming or CDC pipelines from scratch.
- Bring production experience with vector databases and embedding workflows.
- Be highly proficient in SQL and Python.
- Know Snowflake, BigQuery, or Databricks and dbt.
- Have contributed materially to an end-to-end production data environment.
- Understand B2B SaaS recurring-revenue data models.
- Work independently across technical and business stakeholders.
Nice to have
- Experience with Fivetran or Airbyte.
- Knowledge of Looker, Tableau, or Power BI.
- Experience using AI-assisted development tools such as Claude Code.
What’s on offer
- Remote work for candidates in the United States.
- Annual base salary of $172,000 to $215,000.
- Potential benefits including health coverage, flexible paid time off, parental leave, fertility and adoption assistance, 401(k), and educational reimbursement.
Questions about this role
Is the Staff Data Engineer role remote?
Yes. The job is remote and the listed location is the United States.
What is the salary range?
The stated annual base salary range is $172,000 to $215,000.
How much data engineering experience is required?
The role requires at least seven years of hands-on data engineering experience.
Do I need vector database experience?
Yes. The posting requires production experience with vector databases and embeddings.
Related roles
Senior Manager, Financial Planning & Analysis (Remote)
Harbor Compliance is hiring a U.S.-remote Senior Manager of FP&A to lead forecasting, budgeting, SaaS KPI reporting, and financial modeling for a private equity-backed SaaS company. It suits an experienced finance professional with subscription-metrics, operating-model, and planning expertise.
Lead Data Engineer – Databricks & AI | Remote Canada
Lead data engineer role for a financial services client, requiring 10+ years of Databricks expertise and experience collaborating with AI/ML teams to build scalable data infrastructure.
Data Analyst - Data Quality & Databricks
A data analyst role centered on data quality, profiling and reconciliation using Databricks, SQL and PySpark, open across Bangalore, Chennai and Pune for candidates with 7-11 years of experience.
Data Engineer, Data Platform & ML (Remote)
A Poland-remote Data Engineer role building a cloud-based corporate data warehouse for business analytics and machine learning. It suits a data engineer with at least three years of experience in Python, SQL, warehouse design, and modern data tooling.
Data Engineer, Data Platform & Machine Learning (Remote)
A Serbia-remote data engineering role building Tabby's corporate data warehouse, data integrations, and pipelines. It is intended for a data engineer with warehouse design, Python, SQL, and cloud data-stack experience.
Senior Data Engineer, Data Platform & ML (Remote)
A remote senior data engineering role in Yerevan focused on feature-store development, data services, streaming, and production data pipelines. It is suited to an engineer with scalable systems experience across data, machine learning, or backend development.