Senior Analytics Engineer
Valley Stream, Nassau County, New York, 10261, USA
Listed on 2026-07-21
-
IT/Tech
Data Engineering, Data Analyst
The Research & Intelligence team builds and maintains the analytics data platform that powers Carbon Arc’s products — from raw data ingestion and entity mapping to certified data pipelines, knowledge graphs, and AI-powered applications. We own the full lifecycle of alternative data: cleaning, transforming, enriching, scoring, and serving it to clients through APIs, dashboards, and intelligent search.
We are a high-trust, high-impact team that takes end-to-end ownership of complex data problems, working across petabyte-scale datasets in a fast-paced, collaborative environment.
We’re looking for a Senior Analytics Engineer to own and advance critical data pipelines, build production-grade analytics infrastructure, and contribute to the ML and AI products that differentiate Carbon Arc. You will work across the full data stack — from raw extraction through certified delivery — designing scalable systems that ensure data quality, enrich our ontology, and power knowledge graph-driven applications.
This role blends data engineering rigor with analytical depth and product sensibility.
- • Own one or more alternative datasets end-to-end through our data pipeline (extraction, cleaning, transformation, certification, and delivery), ensuring data quality and timeliness across petabyte-scale workloads.
- • Build and maintain Entity Explorer Data (EXD) tables that serve as the foundation for client-facing APIs, dashboards, and analytical products, including demographic, financial, and behavioral metrics at varying entity and temporal granularities.
- • Design and execute ontology mapping pipelines — resolving entities (brands, companies, products, locations) across disparate data sources using deterministic and probabilistic methods, and maintaining mapping infrastructure at scale.
- • Develop and run automated EDA and confidence scoring pipelines to evaluate data quality, detect anomalies, and quantify panel representativeness across datasets and time periods.
- • Contribute to the construction and maintenance of the Carbon Arc knowledge graphs.
- • Build and extend internal tooling and infrastructure - including report templates, monitoring frameworks, bulk data generation, and Pydantic-based validation models - to improve team velocity and pipeline reliability.
- • Collaborate with Engineering, Product, and Insights teams to translate business requirements into scalable data solutions, and participate in cross-functional initiatives such as data infrastructure migration, metadata management, and platform cost optimization.
- • 4+ years of experience building and maintaining production data pipelines and analytics infrastructure at scale.
- • BS or MS in Computer Science, Data Science, Statistics, Engineering, or a related quantitative discipline.
- • Deep proficiency in Python and SQL, with hands-on experience writing PySpark and working with distributed query engines (Trino, Spark).
- • Strong experience with columnar data formats and lakehouse architectures (Apache Iceberg, Parquet, S3-backed data lakes).
- • Familiarity with data transformation frameworks such as DBT and pipeline orchestration tools (Airflow, or equivalent).
- • Experience with entity resolution, ontology design, or semantic data modeling across structured and semi-structured datasets.
- • Demonstrated ability to take ambiguous data problems from scoping through production deployment with minimal oversight.
- • Clear and precise technical communication, with a track record of strong documentation and cross-functional collaboration.
- • Hands-on experience with graph databases (Neo4j) and knowledge graph construction, including ontology-driven node/relationship modeling.
- • Familiarity with GenAI tooling and LLM-based applications, including RAG architectures, prompt engineering, and tool-use frameworks (e.g., MCP).
- • Experience with Star Rocks, Click House, or other OLAP engines for high-performance analytical queries.
- • Experience building internal developer tools, CLI applications, or data quality monitoring dashboards.
- • Fully paid healthcare benefits (Medical, Dental, Vision)
- • Remote work options
- • Paid time off
- • Generous…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).