Category: Data Science
-
I Let AI Agents Play a 5-Year Beer Game: Why Sharing Data Beat Buying a Smarter Model
I ran the classic Beer Game for 260 weeks with AI agents in the chairs. Giving the chain real customer demand cut in-game cost by up to 40 percent, far more than paying for a smarter model did.
-
GLM-5.2 and the Benchmark Trap: How One Score Becomes Two Headlines
A Chinese open-weight model just topped the leaderboards and undercut GPT-5.5 on price by 7x. Then the benchmark tables started disagreeing with each other.
-
Sole Source: The $900k Median Problem Your Dual-Source Checkbox Won’t Fix
The dual-source flag teams buy to feel safe moves the median cost of a disruption about $44k. In the wrong direction. The dependency nobody flags moves it $474k.
-
The Resilience Ladder: Why the Things You Buy to Feel Safe Don’t Save You
I pulled 3,000 disruptions to find what separates firms that survive a shock from firms that bleed. The dual-source checkbox wasn’t it.
-
I Gave an AI Agent the Reorder Button: It Rebuilt the Bullwhip in 250 Days
An AI agent with the reorder button hit ~100% fill rate and looked like a star. Then I measured what it dumped on its suppliers.
-
Global Forecasting with XGBoost in R: A Walmart Weekly Walkthrough
A hands-on walkthrough of global XGBoost forecasting in R with tidymodels and modeltime, applied to the Walmart weekly sales dataset. What the feature importance reveals, when ML earns its complexity, and when ETS or SNAIVE quietly wins.
-
I Ran 6 Models on Real Demand Data — Here’s How I Picked the Winner
Six forecasting models, one real demand series, one honest horse race. Here’s the model that won — and the metric that made the choice unambiguous.
-
Is Your Forecast Any Good? The Forecaster’s Toolbox
Four acronyms decide whether you trust a forecast: MAE, MAPE, RMSE, MASE. Here is when each one lies to you — and the one benchmark that catches them all.
