Years of Experience 2 + Years of Experience
ML experience is a must — feature engineering, model
evaluation, and hands-on application of data-science
techniques, not just pipeline plumbing.
Working knowledge of statistical record-linkage methods such
as the Fellegi–Sunter model: estimating field-level match/non-
match weights from labelled data, likelihood-ratio based scoring,
and setting evidence-driven decision thresholds.
• Proven experience with large-scale ETL: ingesting
heterogeneous databases, mapping them to a common data
model, and profiling/repairing data quality (placeholders, junk
values, inconsistent encodings).
• Strong Python (Pandas / PySpark or similar), SQL /
Postgres, and Parquet-based pipelines.
Fuzzy and phonetic string matching, name normalization,
blocking/candidate-generation strategies, and scoring/threshold
calibration.