Skip to content
← Back to job listings

Senior Data Scientist (Big Data R&D, Identity Graph & KYC)

Socure · Miami, United States

External listingfull-time5 days ago

About The Role

Join Socure, a leading provider of digital identity verification and fraud prevention solutions. As a Senior Data Scientist, you will lead the design and deployment of advanced machine learning and graph algorithms on large-scale PII datasets. You will own end-to-end projects, define standards for feature engineering and data quality, and partner closely with product managers and engineers. This is a fully remote position with a comprehensive benefits package.

  • Lead the design and deployment of advanced machine learning and graph algorithms on large-scale PII datasets, owning end-to-end projects from problem definition through production validation.
  • Architect and optimize graph-based identity representations to improve match rates, reduce false positives/negatives, and support downstream fraud and KYC models.
  • Build and maintain scalable data pipelines and feature stores in Spark/PySpark (or Scala), including data normalization, deduplication, and feature computation across large PII datasets.
  • Demonstrated ability to lead medium‑to‑large projects end‑to‑end, make sound trade‑off decisions under ambiguity, and influence cross‑functional stakeholders with data and clear reasoning
  • Experience in at least one of: identity verification, fraud detection, credit risk, or adjacent high‑stakes domains is a plus
  • Practical familiarity with graph databases and/or graph frameworks (Neo4j, AWS Neptune, GraphFrames, DGL, PyTorch Geometric) and graph algorithms for clustering, link prediction, and community detection is strongly preferred
  • Experience developing production-quality data pipelines and automated workflows using Airflow or similar orchestration tools
  • Solid SQL skills and experience working with large-scale analytical data stores
  • Strong proficiency in Python (preferred) or Scala, including experience with ML libraries such as scikit‑learn, XGBoost, TensorFlow or PyTorch
  • Deep understanding of supervised and unsupervised learning, feature engineering, model evaluation, and experiment design (A/B testing, holdout strategies, stratification)
  • Master’s degree with 3+ years of relevant industry experience, or Ph.D. with 1+ years of experience in applied ML / data science roles; background in Computer Science, Statistics, Mathematics, or related quantitative fields preferred
  • Extensive experience with Spark or PySpark and distributed data systems (e.g., AWS EMR, Databricks) working on very large, messy datasets

This is an external listing. JobSpring does not represent or verify the employer. Report this listing