Matthew Nelson
AI-native Data Platform Engineer · Berkeley MIDS · P.Eng
Calgary, Alberta · matthewpeternelson@gmail.com · +1-403-667-3022 · https://mnelson.ca · Github: https://github.com/matthewpnelson · LinkedIn: https://linkedin.com/in/matthewpeternelson
Data platform engineer and AI-native builder — 10 years in data, 16 in engineering, Berkeley MIDS and P.Eng. Build production data platforms on Dagster, Snowflake, Polars, DuckDB and dbt, then ship AI-native systems on top of them: LLM-in-the-loop pipelines, agentic developer tooling with Claude Code and MCP, and structured-output services gated by evals. Rebuilt StackDX's US data platform as lead engineer and now run roadmap and code review for the 6-person team, still shipping weekly.
Selected Impact
- Architected a production Dagster platform at StackDX across 15 US states and 945 transformation assets, distilling a 3.8 TB source layer into a 10.3 GB exposure layer serving 300M+ records to customer endpoints.
- Founded twochannel.ai: an AI-native catalog where LLM-in-the-loop pipelines and a custom Claude Code toolchain maintain 18k+ products across 2k+ brands — one engineer, end to end.
- Owned Canlin's regulated annual reserves evaluation for 5 years — the externally-audited valuation feeding the financial statements; technical revisions added +$730M (NPV10).
- Built the dbt Cloud / Snowflake platform that became Canlin's foundation for operational data, BI, and ML; automation from it saved 50+ hours/week across Operations.
- Set roadmap and review code for a 6-person, two-country data team at StackDX, while still shipping weekly in the platform codebases.
Experience
Product Manager, Stack Maps & Public Data · Stack Technologies Ltd.
May 2026 — Presentwww.stackdx.com/
Moved from lead engineer into product ownership after rebuilding the US data platform: roadmap and code review for a 6-person team across two countries, while still shipping weekly in the Dagster and Stack Maps codebases.
Product ownership and team leadership
- Run roadmap, prioritization, and code review for a 6-person team across US and Canadian data codebases — while still shipping in the Dagster platform weekly.
- Built PM tooling as software, in daily use: a Jira roadmap auto-scheduler, human-gated production-failure triage, and skills for weekly updates and Jira templating.
- Set Stack Maps product direction (frontend + API), scoping and prioritizing with the frontend and API engineers against the data-platform realities underneath.
Data Platform Engineer · Stack Technologies Ltd.
Jun 2025 — May 2026www.stackdx.com/
Lead engineer on the USA data platform rebuild — replacing a legacy Python ETL stack with a production-grade Dagster-native architecture on AWS, with Polars and DuckDB doing the heavy transformation work. The platform distills a 3.8 TB source layer of 1.03M regulator files into a 10.3 GB exposure layer serving customer endpoints — 127M well production records, 172M lease production records, and 4.1M wellheaders. Built AI-first as well, on custom Claude Code skills, sub-agents and hooks covering scaffolding, data inspection, formatting and review.
Modern Data Platform Architecture
- Architected a production Dagster platform across 15 US states and 945 assets, distilling a 3.8 TB source layer into a 10.3 GB exposure layer serving 300M+ records.
- Converted the Texas production fact build — the largest on the platform — to a streaming Polars plan, eliminating a repeat production out-of-memory failure.
- 146 contract-driven freshness checks across all 15 states, fanning out to 468 evaluations, threshold-calibrated to kill false-positive alerts.
- Designed reusable exposure/coalesce patterns so legacy mart schemas ride on the new fact-table architecture and new states extend the pattern instead of needing bespoke builds.
AI-native Developer Workflow
- Built a project-specific Claude Code toolchain — a parquet-inspector sub-agent, an auto-format hook, and task skills for pipeline triage and schema work — over a token-optimization layer (618M tokens saved to date).
- Profiled my own Claude Code conversation history to find the highest-frequency ad-hoc work, then promoted the dominant pattern into a first-class sub-agent.
Cloud Infrastructure & DevOps
- Deployed the AWS footprint with Terraform IaC across staging and production — RDS, S3, and ECS compute launched and managed per-job by Dagster.
- Implemented CI/CD (GitHub Actions, ruff, pyright, asset schema-contract tests) plus the branching and deployment automation behind staging → prod releases.
Framework & Team Enablement
- Implemented 10 of 15 state pipelines solo, then onboarded two non-platform engineers who shipped the rest by following the documented conventions and schema contracts.
Founder & Principal Engineer · Twochannel
Jan 2021 — Presenttwochannel.ai
Designed and built Twochannel end-to-end — catalog, ingestion pipelines, search, recommendations, and frontend — rebuilt three times since 2021 as the stack evolved, publicly launched in 2026. The 18k+ product / 2k+ brand catalog is maintained almost entirely through AI-native workflows: LLM-in-the-loop scrapers, structured-output enrichment, semantic dedup, and a Claude Code toolchain operating against the full monorepo. The whole thing is the size that would normally need a small data team behind it, and it's just me.
AI-assisted product catalog at scale
- Brand-scraper → LLM normalizer → reviewed dedup pipeline now maintains 18k+ products across 2k+ brands in Directus with minimal manual intervention.
- A 12-agent Claude Code curation team with human approval gates runs catalog curation — three pipeline defects caught at the gates, zero silent corruptions of production state.
- Wizard captures budget, room and genres; an async build job then returns complete system variants, constrained by signal-chain adjacency and per-category budget allocation.
- Replaced a three-tier dedup cascade and a fixed 0.9 confidence cutoff with one Claude decision service holding customer catalog writes to the curation team's standard.
Platform Architecture
- Turborepo monorepo shipping Next.js (Vercel), Directus (Railway), Algolia search, Neo4j relationships, and Doppler-managed secrets across three hosting targets.
- Shipped GDPR Article 17 erasure — a twelve-step deletion pipeline with audit trail, CI-enforced FK rules, and weekly orphan sweep — backed by a 65-test Playwright E2E suite green on prod.
Senior Data Analytics Engineer · Paramount Resources Limited
Apr 2024 — May 2025paramountres.com
Senior member of the Data Analytics and Integration team — integrating best practices across the full data stack, mentoring ICs, and advocating for a modern data platform. Delivered polished, robust solutions spanning custom web apps, modern data-stack PoCs, data engineering pipelines, and high-fidelity dashboards. Toolkit: Databricks, PowerBI, Spotfire, dbt Core, cube.js, DuckDB; SQL, Python, R, DAX.
Dashboards & Corporate Reporting
- Built the corporate PowerBI dashboard suite (Netback, Operating Cost, Capital) — corporate-to-well drill-downs used by executives, operations, production engineering, and development.
- Automated the weekly well production report end-to-end (Spotfire + Power Automate), removing a half-day weekly manual process from the analytics manager's plate.
Custom Web Applications
- Replaced a manual 2–3-hour Excel frac-scheduling process with a Streamlit optimization app — its sequence plans ran ~8 multi-well pads.
Data Governance & Platform Advocacy
- Local modern data stack PoC (Python / dbt / cube.js / DuckDB / PowerBI in Docker) replacing legacy R + CSV processes.
- Data-governance-committee advocate for cloud-first analytics; educated senior leadership on modern stack design.
Data & Advanced Analytics Lead · Canlin Energy Corporation
Feb 2022 — Apr 2024canlinenergy.com/
Designed, built, and ran Canlin's modern data stack (dbt Cloud, Snowflake, DataRobot, Tableau, Streamlit) and led the Integrated Remote Operating Centre (IROC) analytics. Owned descriptive and predictive solutions end-to-end, mentored junior professionals into data roles, and made the business case to senior leadership for sustained investment in analytics.
Corporate Data Warehouse on dbt Cloud + Snowflake
- Built the dbt Cloud / Snowflake platform that became Canlin's foundation for operational data, BI, and ML — SCADA to Accounting in a single source of truth.
- Automated the weekly production reporting process end-to-end — backend workflows, input sheets, and the report itself — saving 50+ hours/week across Operations.
Data Science & ML Ops
- Deployed a modified XGBoost forecaster inferring daily across 8,000+ wells (30-day horizon), plus 2-year forecasts of the full well set, feeding corporate reporting.
- Built the real-time well-status analytics behind Canlin's remote operations centre (~500 wells), with networkx-based outage impact auto-assessment and anomaly-detection reporting.
- Streamlit SME-labeling app + classification model ranking well risk across the portfolio, with explanations.
- Home-grown ML-ops layer (Streamlit + Snowflake + NocoDB) for labeling and data input, with DataRobot for productionized models.
Corporate Data Scientist & Reserves Manager · Canlin Energy Corporation
Oct 2017 — Jan 2022canlinenergy.com/
Drove Canlin's transition to a self-service data model by pairing data engineering with petroleum-data expertise — implementing Tableau as the company-wide reporting tool and data mart, advocating for repeatable BI, advanced analytics, and automation. Owned the regulated annual reserves evaluation — the externally-audited, company-wide valuation of every corporate asset, tied directly to the financial statements — improving accuracy and unlocking +$730M in value through technical revisions.
Tableau data-warehouse implementation
- Implemented Tableau as the corporate reporting platform — 100+ data sources, 50 workbooks, and 20 Prep flows published in year one.
- Mentored every Tableau Creator/Explorer in the company and made the executive case for self-service data.
Corporate Reserves — regulated annual asset valuation, 5 years
- Owned the externally-audited NI 51-101 reserves evaluation for 5 years — technical revisions added +$730M (NPV10) across four consecutive cycles.
Corporate Acquisitions & Divestitures
- Built the data flows and dashboards for rapid A&D evaluation through a depressed-price divestiture cycle that brought corporate debt from $120M to $0
Exploitation / Development Engineer, Foothills & South · Centrica Energy Canada
Jul 2014 — Sep 2017www.centrica.com/
Exploitation/development engineering across multiple assets. Authored multi-scenario asset-longevity analysis that set 5–10-year corporate strategic direction.
Asset-longevity analysis & internal tooling
- Multi-scenario asset-longevity modeling across all operated facilities that set 5–10-year corporate strategic direction
- Built a VBA-driven opportunity-tracking tool with a staging/approval workflow for development planning
Production & Exploitation Engineer · Pengrowth Energy
Aug 2010 — Jun 2014On-site production engineering plus exploitation engineering across multiple assets. Built multi-dimensional tracking models from time-series sensor data and designed data-driven drill programs.
Highlights across both roles
- On-site production/operations engineering, including long-term sensor-data trending and equipment monitoring
- Multi-dimensional tracking models built from time-series sensor (thermocouple) data
- Drill-program design and feasibility analysis across multiple assets
Engineering Co-op Student · Pengrowth Energy (Co-op)
2006 — 2010Multiple co-op terms as part of the University of Waterloo Engineering program — 8 months at Lanmark Engineering, 4 months at the Olds Sour Gas Plant, and 12 months in the Pengrowth office as an Exploitation Engineering co-op.
Selected Projects
twochannel.ai · twochannel.ai
2021 — CurrentAI-assisted HiFi product catalog & intelligent stereo-system designer. Founder project; see the Twochannel entry under Experience for the full detail. 18k+ products across 2k+ brands maintained through LLM-augmented ingestion, semantic dedup, and a Claude-Code-first developer workflow. Stack: Next.js, Directus, Algolia, Neo4j, Turborepo.
- AI-native monorepo developed primarily through Claude Code (skills, sub-agents, hooks, MCP)
- LLM-in-the-loop ingestion + human-reviewed semantic dedup
- Next.js frontend on Vercel, Directus CMS on Railway, Algolia search, Neo4j relationships
- Doppler single-source-of-truth secret management across Vercel / Railway / GitHub Actions
- 65-test Playwright E2E suite, green on production
StackDX USA Data Platform · www.stackdx.com
2025 — CurrentProduction-grade USA oil & gas data platform at StackDX. Dagster-native orchestration and transformations, Polars/DuckDB performance layer, Terraform-managed AWS infra. Developed AI-first with a purpose-built Claude Code toolchain (parquet-inspector sub-agent, /inspect-parquet + /explore-data + /scaffold-layer commands, auto-format hook).
- 15 US states, 945 Dagster transformation assets, 300M+ served records
- Targeted Polars + DuckDB rewrites and incremental materializations cutting processing time on critical datasets
- Reusable exposure/coalesce-sources patterns for legacy-schema compatibility
- AWS infra-as-code (RDS, ECS, S3) with Terraform
- Custom Claude Code skills, agents, and hooks driving developer velocity
Instill · instillmeditation.ca
2017 — 2018A modern and refined meditation and lifestyle brand. The site remains active; the bulk of the build ran 2017–2018 as a full-stack React app for online course content.
Vedic Meditation Directory · learnvedicmeditation.co
2018 — 2019Helped students find their local teacher of Vedic Meditation. Ran through multiple versions — from a full-stack React app down to the final iteration: a super-simple Next.js static site powered by a single Google Doc so teachers could self-manage listings.
Skills
AI Engineering & Agentic Development: Claude Code (agents, skills, hooks, slash commands, sub-agents), MCP servers & clients, Anthropic / OpenAI / Perplexity APIs, Vercel AI SDK, Structured outputs & schema-adherent LLM calls, Prompt caching & evals, RAG, LLM-in-the-loop data pipelines, Editors: Claude Code, Cursor, Windsurf
Data Engineering & Modern Data Stack: dbt (Core & Cloud), Dagster, Snowflake, Snowpark, Polars, DuckDB, Tableau Prep, cube.js & dbt semantic layers, pandas, SQL, Schema contracts
Cloud, Infra & DevOps: AWS (ECS, RDS, S3), Terraform, Docker, GitHub Actions CI/CD, Vercel, Railway, Supabase, Doppler, pre-commit / ruff / pyright
Data Visualization & BI: Tableau, PowerBI, Spotfire, Plotly, D3.js, matplotlib, Seaborn, Dash, Bokeh, ggplot2
Machine Learning & NLP: scikit-learn, XGBoost, PyTorch, Hugging Face, nltk, Time-series forecasting, Classification, Anomaly detection
Databases: Snowflake, PostgreSQL, DuckDB, Neo4j, networkX, MongoDB, SQL Server, Oracle
Web App Development: Next.js, React, Astro, Directus (headless CMS), Algolia, Node.js, Express, Tailwind CSS, Playwright, Streamlit, Flask, Dash, Turborepo / npm workspaces
ML Ops: Snowflake & Snowpark, DataRobot, Streamlit labeling apps, Active learning workflows
Programming Languages: Python, SQL, TypeScript, JavaScript, Bash, R
Upstream Oil & Gas: Exploitation / Development Engineering, Production Engineering, Operations, Corporate reporting, A&D evaluation
Petroleum Reserves Evaluation: COGEH, NI 51-101
Education
Master of Information and Data Science · University of California, Berkeley
May 2016 — Jun 2018 · GPA 3.888Bachelor of Applied Science, Chemical Engineering · University of Waterloo
Sep 2005 — Jun 2010Certifications
- Professional Engineering Designation (P.Eng) — Association of Professional Engineers and Geoscientists of Alberta (APEGA) (2014)
Awards
- Randy Duxbury Memorial Award — University of Waterloo (2010)
- Ontario International Education Opportunity Scholarship — University of Waterloo (2007)