×
Register Here to Apply for Jobs or Post Jobs. X

Data Engineer - Onboarding North America

Remote / Online - Candidates ideally in
Brampton, Ontario, C6S, Canada
Listing for: Sardine
Remote/Work from Home position
Listed on 2026-08-05
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, Data Engineering, AI Engineer (Applied/Software), Python
Job Description & How to Apply Below

Who we are:

Sardine is the leading agentic risk platform for fighting financial crime. Our integrated solution unifies data across risk teams to help organizations stop fraud in real time, prevent AI-driven attacks, and automate fraud and AML operations. Sardine’s platform is strengthened by one of the fastest-growing fraud consortiums in the market, spanning more than 6 billion profiled devices, 800 million consumers, and 3 million businesses worldwide.

Leading companies including FIS, GoDaddy, Intuit, Edward Jones, Zoom Info, and  rely on Sardine to secure and grow trust in their products.

Our culture:
  • We have hubs in the Bay Area, NYC, Austin, Toronto, and São Paulo. However, we maintain a remote-first work culture. #Work From Anywhere

  • We hire talented, self-motivated individuals with extreme ownership and high growth orientation.

  • We value performance and not hours worked. We believe you shouldn't have to miss your family dinner, your kid's school play, friends get-together, or doctor's appointments for the sake of adhering to an arbitrary work schedule.

Location:
  • Remote - United States or Canada

  • From Home / Beach / Mountain / Cafe / Anywhere!

  • We are a remote-first company with a globally distributed team. You can find your productive zone and work from there.

About the role

We are looking for a Senior Data/ML Engineer to own the data and machine learning foundation that Sardine's compliance decisions run on. Every onboarding decision we make — a payment approved, an account blocked, a KYC case escalated — is the output of a pipeline someone built. This role owns those pipelines end to end: how data arrives, how it becomes a feature, how that feature becomes a model, and how that model stays correct in production.

This is a high-impact, highly technical IC role sitting at the intersection of data engineering and ML engineering. We need someone at the senior level to set technical direction for the next order of magnitude: new feature generation, build specific models around KYC onboarding, in house entity matcher for the sanctions and more

You will write production code, make architectural calls that outlive your tenure, and raise the bar for how a small team ships fraud ML. You will work directly with data scientists, backend engineers, and the fraud analysts who use what you build.

What you'll be doing
  • Own the data ingestion layer that brings device telemetry, transaction events, KYC/identity signals, and third-party enrichment into the platform — designing streaming pipelines (Pub/Sub, Apache Beam on Dataflow, Flink) and batch pipelines (Python, Airflow on Cloud Composer, Spark on Dataproc) that are correct, observable, and cheap to extend.

  • Build and evolve our feature platform, where the same Chronon feature definitions are computed by Flink for streaming and Spark for batch, with aggregation windows from one hour to 300 days, served to the rules engine and to models under a sub-second budget.

  • Establish feature correctness as an engineering discipline: streaming-versus-batch reconciliation, recomputation tests against the warehouse, train/serve parity checks, and drift monitoring that catches a broken feature before an analyst does.

  • Productionize fraud and identity ML models — training pipelines on Vertex AI and Kubeflow, gradient-boosted and tree-based models (XGBoost, LightGBM, Cat Boost, scikit-learn), hyperparameter search, SHAP-based explanations, and score normalization — and build the automated retraining, champion/challenger promotion, and rollback machinery we don't yet have.

  • Engineer KYC, AML, and identity risk signals: document verification and doc-KYC outcomes, sanctions/PEP/adverse-media screening results, email and phone risk, synthetic identity indicators, bank and account verification, and periodic customer due diligence — turning noisy, multi-vendor, multi-jurisdiction data into features a model can actually learn from.

  • Integrate and harden new data sources, including 30+ third‑party enrichment providers called in parallel on the request path, plus our cross‑client consortium network — owning failover behavior, timeout budgets, graceful degradation, caching, and cost.

  • Own the…

Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary