Explorer, AI-Native Advocate, Behavior Designer
Leon is a builder and an AI-native advocate. He develops and evaluates AI products, measurement systems, and operating practices designed to turn isolated AI experiments into useful, repeatable work. His approach combines behavior design, data science leadership, causal inference, and machine learning.
At Dropbox, Leon works on Dash AI, where he builds evaluation systems that connect model and product quality to user outcomes, latency, cost, and product value. He also designs reusable AI workflows, shapes adoption strategy, and rigorously validates decision models so product decisions can be evaluated before and after launch.
Earlier, Leon developed quantitative solutions for unstructured enterprise problems at Slack, led data science for advertising product prototypes at TikTok, and used experimentation to tackle difficult earner problems at Uber.
Current focus
- AI-native adoption: Designing the workflows, incentives, review practices, and cultural mechanisms that help teams use AI well.
- AI product evaluation and quality: Defining model and product quality, linking it to user behavior, and measuring cost and latency alongside value.
- Decision science: Pressure-testing models, separating association from causality, and making uncertainty visible before decisions are made.
Skills
Methods and systems Leon has used to build, measure, and improve products.
Languages and Infrastructure
Python (PyData, SciPy, TensorFlow, PyTorch) • R (dplyr, ggplot2, shiny) • SQL (Presto, Spark SQL, PostgreSQL) • JavaScript • Spark, Flink • Airflow, Docker, Kubernetes
Data Science and ML
AI Product Measurement • Causal Inference • A/B Testing and Quasi-Experimental Methods • LLM Evaluation and Prompt Engineering • RAG and Retrieval Optimization • Classification, Regression, Structured Prediction • Time Series Analysis • Multi-Armed Bandits
Domain
AI-Native Adoption • Enterprise AI Products • Enterprise Level Experiments (ELE) • Behavioral Design • Marketplace Optimization • Advertising Platforms
Experience
Dropbox Staff Data Scientist, Tech Lead · Dash AI 2026 - Present
- Builds measurement frameworks for AI product experiences across quality, engagement, latency, cost, and experimentation.
- Designs AI-native workflow practices to support reuse, clear ownership, human review, and governance.
- Brings quantitative insight into product design early by shaping metric trees, instrumentation, and decision criteria before launch.
Slack Staff Data Scientist 2024 - 2026 · Sr. Data Scientist 2022 - 2024
- Led the design and refinement of Slack AI’s data organization and infrastructure, managing a cross-functional team across experimentation, data quality, and model development.
- Partnered with product and engineering teams to define success metrics and evaluation frameworks for generative AI experiences, while improving the underlying analytical data foundation.
- Developed a multi-signal model for enterprise workspace health and introduced Enterprise Level Experiments to support product decisions at the organization level.
TikTok Senior Data Scientist 2021 - 2022
- Led data science for advertising product prototypes and automated ad solutions, including north-star metric design, impact estimates, and value analysis.
- Designed and advised on experiments for product launches, translating results into product and platform decisions.
- Developed experimentation guidance and an organizational runbook covering test selection, low-sample settings, pre-existing bias, implementation, and analysis.
Uber Data Scientist 2019 - 2021
- Designed and executed A/B tests, difference-in-differences studies, and switchback experiments for new earner products, translating results into product decisions.
- Quantified addressable markets, forecasted product goals, and contributed evidence to roadmap development.
- Led instrumentation standards, owned end-to-end analytical data workflows, and improved a logistic-regression search system as product needs evolved.
OfferUp Product Data Science 2017 - 2019
- Helped design and implement an in-house experimentation platform with randomized bucketing and multiple statistical tests, substantially reducing analysis and deployment time for revenue features.
- Diagnosed non-orthogonal experiment assignment by tracing unexpected behavior to the hashing approach and confirming the relationship with a chi-squared test.
- Built KPI, analysis, and visualization pipelines for new revenue streams, and used experimentation to shorten personalization-model iteration time while increasing advertising revenue.