JM.Download CV
Data scientist + financial economist

Turn messy data into decisions you can defend.

I use Python, SQL and statistical modelling to answer commercial and research questions across product analytics, risk, energy markets and applied AI.

UK-based · Open to relocation to the UAE · One-month notice

Jacob Mackey outdoors in a narrow city street
Jacob MackeyUK → UAE
Event-level analytics2.76mretail events analysed
Time-series validation8,615out-of-sample hours
Risk modelling0.8075held-out PR-AUC
Research disciplineNo edge claimed.when uncertainty crossed zero
Profile01 / 05

Financial economics meets applied data science.

I work from the decision backwards. My postgraduate training gave me a grounding in markets, econometrics and uncertainty; my project work applies that grounding to product behaviour, energy forecasting, fraud risk and evidence-grounded AI.

I prefer questions where the data is imperfect and the answer must change a real choice. I acquire, clean and test data, compare models with credible baselines and make weak or null results visible. A model that should not be deployed is still useful if it prevents a costly decision.

I am targeting data science, analytics, risk and quantitative research opportunities in the UAE, where I want to build long-term experience in teams making high-consequence decisions.

Jacob Mackey working at a desk
Behind the analysis
A working principle
“The useful answer is not the most impressive model. It is the one that survives scrutiny and changes the decision.”
QuestionEvidenceDecision
Selected work02 / 05

Five questions. Five defensible decisions.

Forecasting, product analytics, information retrieval, agentic research and risk modelling—connected by the same standard of evidence.

01SQL · Product analytics

Retail product journey

Where do users leave the path from product view to basket and transaction, and what happens after activation?

I modelled an ordered funnel from 2,756,101 events, then extended it to first-transaction activation, weekly cohorts and event sequences in DuckDB. The result identifies the measurement priority without pretending the dataset can prove a marketing intervention worked.

SQLDuckDBFunnelsCohorts

Decision: diagnose the view-to-basket step and collect commercial context before choosing an intervention.

Read the evidence

What I built

I wrote a reproducible workflow for event profiling, ordered funnels, first-transaction activation, weekly cohorts and event transitions, with compact CSV and JSON evidence.

Validation

Funnel stages must occur in order for the same visitor. Committed headline JSON and funnel CSV outputs agree on the reported values.

What I found

Ordered conversion was 2.35% from view to basket, 30.22% from basket to transaction and 0.71% end to end. The largest measured loss occurs before basket addition.

Boundary

The data do not contain campaign exposure, margin, acquisition cost or experiment assignment, so they cannot establish promotion effectiveness, incentive profitability or causal impact.

View repository
Ordered journey
View100%
Basket2.35%
Transaction0.71%
02Risk · Classification

Fraud alert thresholds

How should rare fraudulent transactions be identified while balancing missed fraud, false positives and review capacity?

I built a reproducible workflow around the real 284,807-row public dataset using chronological train, validation and test splits. A transparent logistic model became the operating baseline, with its threshold treated as a policy choice rather than a leaderboard score.

PythonLogistic regressionPR-AUCThreshold policy

Decision: choose the final threshold with fraud operations, using capacity and customer-friction tolerances.

Read the evidence

What I built

I built data validation, chronological splits, logistic baselines, a boosted-tree challenger, validation-only threshold selection and held-out reporting. The model artifact records preprocessing, weights, threshold and metrics.

Validation

The full run contains 284,807 transactions and 492 fraud cases. I selected the model and threshold on validation, then used the test set once.

What I found

The selected model returned 82.61% precision, 76.00% recall and 0.8075 test PR-AUC. Precision, recall and alert volume are more informative than accuracy for this imbalanced problem.

Boundary

The short, anonymised sample does not establish production stability. PCA components are debugging features, not customer-facing reason codes, and the cost weights are illustrative.

View repository
Held-out test
Precision82.61%
Recall76.00%
PR-AUC0.8075
03Agentic AI · Energy markets

European power research agent

Can an LLM produce better multi-step power-market research when it retains its reasoning and accumulated evidence rather than repeatedly reconstructing the investigation?

Inspired by OpenAI’s ARC-AGI-3 harness research, I built and preregistered a controlled OpenAI Responses API experiment using GPT-5.6 Sol at max reasoning. Eight hash-locked synthetic European power-market investigations each required five sequential evidence gates. Every case was run twice across three conditions: stateless with a three-action rolling history, retained reasoning, and retained reasoning with deliberately forced compaction—48 randomized API runs in total.

Retained reasoning achieved 16/16 exact successes and a 100/100 mean score, versus 0/16 and 0/100 for the stateless baseline. The paired episode-level difference was +100 points (95% CI [100, 100], exact sign-flip p=0.0078), while using 35.2% fewer output tokens and costing 41.9% less. Every stateless run exhausted the fixed step budget after losing and reconstructing earlier evidence.

The primary contrast saturated: all eight episode-level score differences were +100. That is why the bootstrap interval collapsed to [100, 100], while p=0.0078 is the minimum attainable two-sided exact sign-flip value with eight non-zero episode differences. The result supports retention in this defined stress test but does not provide fine-grained effect-size resolution.

Forced compaction activated in all 16 designated runs, but it did not meet the preregistered quality-preservation or output-token-reduction rules. It scored 97.5/100 with 14/16 exact successes, used 8.0% more output tokens, cost 39.2% more and took 198.6% longer than retained reasoning alone.

PythonGPT-5.6 SolResponses APIPreregistered evaluation

Decision: The experiment supports retained reasoning for this defined synthetic memory-stress workflow. It does not support a compaction benefit at the aggressive 1,000-token threshold, and it does not claim live-market alpha or universal generalisation.

Read the evidence

What I built

I preregistered eight hash-locked synthetic investigations with five sequential evidence gates, then ran every case twice across three randomized harness conditions using the same GPT-5.6 Sol model, max reasoning, tools, evaluator and fixed eight-step budget: 48 API runs.

Validation

The paired analysis treated the eight episodes—not the repeated draws—as the independent units. All eight episode-level score differences were +100, so the episode-clustered bootstrap interval collapsed to [100, 100]. The exact two-sided sign-flip p-value of 0.0078 is the minimum attainable with eight non-zero episode differences.

What I found

Retained reasoning scored 100/100 with 16/16 exact successes, compared with 0/100 and 0/16 for the stateless baseline, while using 35.2% fewer output tokens and costing 41.9% less. All stateless runs exhausted the step budget after reconstructing discarded evidence.

Boundary

This deliberate memory-stress benchmark saturated on the primary contrast, so it supports retention for this defined synthetic workflow but cannot finely resolve the effect size or establish generalisation to live research. Forced compaction scored 97.5/100 with 14/16 exact successes, used 8.0% more output tokens, cost 39.2% more and took 198.6% longer than retention alone, and did not meet the preregistered quality-preservation or token-reduction rules.

Official footage · Autoplays muted
Research inspiration: OpenAI’s ARC-AGI-3 harness comparison. Embedded from the original publication.Read OpenAI’s source article
Controlled harness comparison
ARolling history0/100 mean
BRetained reasoning100/100 mean
CReasoning + compaction97.5/100 mean
Held constantModel · tools · sealed episodes · evaluator
Synthetic memory-stress set · Retention supported
04Forecasting · Energy markets

German wind forecasting

Could a residual model improve Germany’s published wind forecast enough to support a credible power-market signal?

I built an XGBoost residual calibrator around SMARD’s public forecast and tested it across 12 expanding walk-forward folds. The pooled gain was small and statistically uncertain, so I reported no reliable incremental edge and froze a stricter day-ahead design.

PythonXGBoostWalk-forward validationNewey-West

Decision: retain the public benchmark and do not claim or deploy an unsupported trading edge.

Read the evidence

What I built

I joined public wind, archived weather and day-ahead price data, added reproducible QA, trained a residual calibrator and saved fold errors, uncertainty diagnostics and a paper-signal test.

Validation

The historical reference run covers 8,615 out-of-sample hours across 12 expanding folds. I compared it with SMARD and used Newey-West and delivery-day bootstrap inference.

What I found

SMARD MAE was 1,389.44 MW and XGBoost MAE was 1,384.67 MW: a 0.34% pooled improvement. Both skill intervals crossed zero, mean fold skill was negative, and the paper signal was not statistically persuasive.

Boundary

The historical run does not satisfy the stricter frozen issue-time information set. Its paper signal also lacks executable entry, settlement, liquidity and cost assumptions.

View repository
Pooled MAE gain
+0.34%
−2.91%CI crosses zero+3.60%
05Applied AI · Information retrieval

Grounded energy research

How can an analyst retrieve answers from long market reports without allowing unsupported synthesis to pass as evidence?

I built a retrieval pipeline over public ENTSO-E and ACER reports with checksummed ingestion, page-aware chunking, local embeddings, constrained generation and an explicit refusal when evidence is absent. The system accelerates evidence discovery while keeping claim-level review human.

PythonChromaDBEmbeddingsResponses API

Decision: use the pipeline for research triage and drafting—not as an autonomous source of truth.

Read the evidence

What I built

I created a pinned source manifest, page-aware chunking, local sentence-transformer embeddings, cosine retrieval in ChromaDB and grounded generation through the OpenAI Responses API.

Validation

Validation cases cover expected-source retrieval, grounded phrases, an absent-evidence refusal and unretrieved filenames. Each source URL and file hash is pinned.

What I found

The pipeline creates a reviewable chain from source manifest to retrieved passages, answer and citation result. The repository supports an implementation claim, not a benchmark-performance claim.

Boundary

A cited filename being present in retrieved context does not prove that every sentence follows from that passage. Prompts and refusal logic target failure modes; none guarantees accuracy.

View repository
Evidence path
01Sources
02Chunks
03Retrieve
04Answer
05Review
Applied work & research03 / 05

Scope made explicit.

Employment, independent research and assessments are labelled so the reader never has to infer what the work was.

01
Self-initiated analysis within Customer Service Associate role

RMS pricing and utilisation

I proposed a management question around how facility rates and utilisation should vary across facilities, day types and time bands for the 2026/27 academic year. With permission, I built Python QA and an Excel decision model. The recommendations were scenario-based and had not been implemented, so I make no realised-revenue claim.

02
Independent research

MSc dissertation

I investigated whether investor overconfidence, proxied by trading volume, was associated with U.S. equity returns. The 5,280-observation panel used IV/2SLS after an endogeneity diagnosis. The final specification found a negative contemporaneous association, but limited model fit and design constraints prevent a causal interpretation.

03
Final-stage technical assessment

Ekimetrics commercial case

I analysed weekly sales evidence, interpreted a supplied regression and presented commercial recommendations. I did not build or rerun the regression, and without margin, cost or out-of-sample evidence I did not claim causal or profit impact.

04
Technical assessment

Revolut financial-crime case

I audited activation and KPI definitions, normalised geography comparisons and built a multi-factor priority queue. The feedback reinforced two habits: define economic activity narrowly and investigate the mechanism beneath high-level behaviour.

How I work04 / 05

Rigour that survives contact with a real decision.

01

Start with the decision

Define who will use the answer, what choice it informs and what would change that choice.

02

Audit the data

Test provenance, timestamps, missingness, duplicates and metric definitions before modelling.

03

Earn complexity

Compare a model with a credible operational or statistical baseline—not with no benchmark.

04

Respect time

Validate chronologically when a random split would let the future leak into the past.

05

Make it usable

Pair the result with uncertainty, its limits and a practical recommendation a stakeholder can challenge.

Technical range05 / 05

Tools in service of the question.

The portfolio spans modelling, product analytics, research and AI—but the method stays consistent: inspectable inputs, reproducible work and a clear decision boundary.

Programming & data

PythonPandasNumPySQLDuckDBExcelREST APIs

Modelling & statistics

XGBoostscikit-learnClassificationTime-series forecastingIV / 2SLSNewey-West inference

Analytics & reporting

FunnelsActivationCohortsData QAScenario analysisStakeholder reporting

Applied AI

OpenAI Responses APIStructured outputsEmbeddingsChromaDBEvidence retentionEvaluation cases

Reproducibility

GitGitHubpytestPinned dependenciesChecksumsCommitted outputs
Education
2024

MSc Financial Economics (Merit)

Cardiff University

2023

BSc (Econ) Economics & Finance

Cardiff University