Senior Data Scientist, Lead Data Scientist
  • NA
3 Days Ago
NA
C2C
Remote
5-10 Years
Required Skills: Python and SQL
Job Description
  • The candidate must be able to name the industry and the outcome variable for each such engagement. Retail same-store analysis is the classic form; the analogue here is comparing similar schools and events rather than following one trend line.
  • Presents to non-statisticians: business outcome first, method second; confidence stated in plain language; explicitly states what the forecast cannot do; never opens with an undefined statistical term.
  • Can teach the method to a client team, not only execute it.
  • Participate actively in stand-ups and backlog refinement, engage business stakeholders directly, understand why the business is asking a question, and challenge or refine the request when it is wrong. 
  • Strategic recommendations are expected alongside hands-on delivery.
 
Qualifications Required: -
 Must be able to work EST hours
  • 5+ years of applied forecasting.
  • Two or more comparable forecasting engagements led start to finish.
  • Comparable-unit / "same-store" forecasting experience.  
  • Executive communication. 
  • Thought leadership. 
  • Multivariable regression, plus collinearity analysis and VIF interpretation.
  • Forecast model development, tuning, selection and holdout validation.
  • Metric fluency: R², WAPE, MAPE, p-values — and why WAPE is used at event grain (many events sell zero, which breaks MAPE).
  • Sparse and zero-inflated data. Many variables populate on under 25% of events, some as low as 10%. Nulls must never be silently treated as zeros.
  • Data-leakage discipline and point-in-time correctness: every feature must exist before the event starts.
  • Python and SQL; reproducible notebooks.
  • Snowflake, including Snowflake ML Model Registry (model versions carry metrics and training-dataset references).
  • Git and pull-request workflow; all code merged to the client repository, no private forks.
Preferred:
  • Architecture Decision Records (ADRs) and written process documentation.
  • Categorical encoding at scale (~30–35 source variables expand to ~70 columns).
  • Sports, streaming, ticketing or subscription-business domain exposure.
  • Hierarchical or mixed-effects models for low-volume segments.

Jobseeker

Looking For Job?
Search Jobs

Recruiter

Are You Recruiting?
Search Candidates