How Up or Down works
Up or Down asks one question of 32,453 daily prices from 5 markets: will tomorrow's close be higher than today's? It trains 7 machine-learning models to answer, trades on their answers in a backtest, and shows every step as an interactive chart. This page explains what it does, how, and what you can do here.
- Markets
- 5
- Daily prices
- 32,453
- Years, 2000 to 2026
- 26
- Features per day
- 20
- Models
- 7
- Latest close · 2026
- 25 Sep
Verdict so far
Tomorrow's direction is close to a coin flip, and every page says so
Across 35 model and ticker pairs, test ROC AUC runs from 0.48 to 0.53, where 0.50 is chance. The model each ticker would pick by cross-validation beats the naive baseline on 0 of 5 tickers. The value here is a sound method and clear charts, not a money machine.
What it does
Three jobs, in order, for every market in the dashboard
-
01
Predict
Each trading day becomes a row of 20 numbers describing recent price action, trend, momentum, volatility and volume. 7 classifiers learn from past rows to answer will tomorrow's close be higher? A second target asks: will tomorrow move more than a typical day of the past year?
-
02
Backtest
A simple strategy follows each model: long when it predicts a rise, short (or in cash) when it predicts a fall, paying 5 basis points every time the position changes. Its returns, risk and drawdowns are set against simply buying and holding.
-
03
Explain
Every stage is an interactive chart with a one-line "how to read" and a caption computed from the data that states the takeaway, including when a model fails. Nothing is cherry-picked: models are compared with naive baselines on days they never saw.
How it works
From raw prices to a verdict in six steps. Each step is a small, tested function in the core library.
-
01
Collect
Daily open, high, low, close and volume come from Yahoo Finance for S&P 500, Amazon, Microsoft, Alphabet (Google) and Oracle, adjusted for splits and dividends, from 2000 (or the company's listing date) to the latest close. Prices are cached as Parquet files, so pages never wait on a download.
stockml/data/loader.py -
02
Clean
Prices are validated, and bad ticks are repaired. A daily move is flagged only if it is more than 6 times the volatility of the previous 63 days and reverses the next day, the signature of a data error. Genuine crashes such as 2008 and March 2020 are kept.
stockml/data/cleaning.py -
03
Engineer features
20 features per day, all expressed as ratios or percentages so they compare across decades and across stocks:
- Price action the day's high–low range, open-to-close change and one-day return
- Trend the close against its 3, 10, 30-day simple and 3, 10, 30-day exponential moving averages, and whether the 10-day EMA sits above the 30-day one
- Momentum 14-day RSI and MACD (12, 26, 9) with its histogram
- Volatility realised volatility over 5, 21 and 63 days, their ratio, and the close's position within 20-day Bollinger Bands
- Volume the day's volume against its 20-day average
The label for day t is 1 if the close on day t+1 is higher. Features for day t use only data up to that day's close, and tests prove it: they change future prices and check that no earlier feature moves.
stockml/features/ -
04
Train
History is split in time, never shuffled: the first 80% of days train the models and the last 20% are held back as the test period. Each model is one scikit-learn pipeline (scaling, optional PCA, then the classifier), so preprocessing is fitted on training folds only. Settings are tuned with 5-fold expanding-window cross-validation inside the training period.
- Logistic Regression tunes C from 0.01, 0.1, 1.0
- SVM (RBF) tunes C from 0.3, 1.0, 3.0
- Linear SVC
- Random Forest tunes max_depth from 3, 6
- Extra Trees
- Bagging (KNN) PCA to 5 components
- Gradient Boosting tunes max_depth from 2, 3, 5
stockml/models/ -
05
Evaluate
The same fitted pipeline is scored on the test period with accuracy, precision, recall, F1 and ROC AUC, next to a naive baseline that always predicts the more common direction. ROC AUC gets a 95% confidence interval, and a walk-forward check refits every model each 63 trading days to see whether fresher data helps. Feature importance is measured on the last training fold, never on the test set.
stockml/evaluation/metrics.py,stockml/models/walk_forward.py -
06
Backtest
The prediction made at day t's close sets the position that earns the log return from t to t+1. Costs are charged on every position change, including the first entry. Sharpe and Sortino are annualised over 252 trading days, max drawdown is measured on the equity curve, and Calmar is annual return divided by max drawdown. A cost sweep from 0 to 50 bp shows how fast turnover eats an edge.
stockml/evaluation/backtest.py
How to read Grey bars are the days a model learns from; blue bars are the cross-validation folds that score it while settings are chosen; orange is the test period (Jun 2021 to Sep 2026), touched once at the end. The walk-forward row refits at the start of each orange block using every earlier day.
Ground rules
What keeps the results honest. Each rule is enforced in code and covered by tests.
No look-ahead
A feature on day t sees nothing after day t's close, and the final day, whose outcome is unknown, is dropped.
Time only moves forward
Chronological train/test splits and expanding-window cross-validation. Shuffled k-fold, which would train on the future, is never used.
Preprocessing inside the pipeline
Scalers and PCA are fitted on each training fold, never on the full history, so no test-period statistics leak in.
One model, everywhere
The pipeline that is tuned is the one that is scored, backtested and plotted. No refitting a friendlier variant for the report.
Choose by the past only
When one model is picked per ticker, the choice uses training-period cross-validation, so its test result stays out-of-sample.
Always show the baseline
Every model sits next to the naive guess and buy & hold, and the page says so when it loses.
What you can do
Six pages, each answering one question. Pick a ticker on any page and it follows you to the next.
-
Overview
How has this market behaved?
Candlesticks with 50- and 200-day moving averages and market events marked (Dot-com peak to ChatGPT launch), volume, the fall from each all-time high, rolling volatility, returns by calendar year, and a year × month heatmap.
Try Set a start and end date to zoom the whole page into one episode, such as 2008 or 2020.
-
Indicators
Do classic trading signals work?
Moving averages, Bollinger Bands, RSI with overbought and oversold zones, and MACD, with 10/30-day EMA crossovers marked. Below, the share of up days after each signal, with 95% intervals.
Try Toggle overlays with the buttons above the chart, then check whether any signal clears the dashed average.
-
Exploration
Is there any signal before a model is trained?
Return autocorrelation, fat tails against a normal curve, each feature's correlation with the next day's return against a noise band, next-day up-rate by feature decile, distributions by outcome, class balance and feature redundancy.
Try Use the dropdown on the decile chart. A useful feature slopes clearly from the lowest to the highest decile.
-
Models
Can a model beat a coin flip?
A comparison table against the naive baseline, cross-validation scores and their fold-by-fold stability, ROC curves, rolling test accuracy, walk-forward results with confidence intervals, and for one chosen model its score separation, confusion matrix and feature importance.
Try Switch "Predict" between direction and volatility, or click a model's name for its detail charts.
-
Backtest
Would trading on the predictions have paid?
Equity curves against buy & hold, a risk table (return, volatility, Sharpe, Sortino, drawdown, Calmar, hit rate, trades and costs), cost sensitivity, risk against return, drawdowns, rolling Sharpe, the daily-return histogram and monthly returns.
Try Drag the cost slider from 0 to 50 bp, or switch to long/flat to sit in cash on predicted down days.
-
Multi-ticker
Does anything work consistently?
A ticker × model heatmap, each ticker's cross-validation pick against buy & hold, rebased prices, risk against return for each stock, and how closely the markets move together.
Try Switch the heatmap between accuracy, ROC AUC and Sharpe ratio.
Working with the charts
Every chart is live, not a picture
- Hover or tap for exact values. Time-series charts show every series for that day together.
- Drag to zoom into a stretch of time; double-click to reset. Range buttons jump to the last 6 months, 1 or 5 years.
- Click a legend entry to hide or show that series; double-click to show it alone.
- Save a chart as a PNG from the camera icon in the toolbar that appears on hover.
- Switch theme with the moon and sun button at the top right. Your choice is remembered.
- Share a view. Settings live in the address bar, so a link such as
/backtest/?cost_bps=20&mode=long_flatreopens exactly that page. - On a phone, chart titles sit above each plot, legends move below it, and wide tables show key columns first, with the rest a tap away.
The data
What is in the dashboard right now
| Market | Symbol | Quoted in | From | To | Daily bars |
|---|---|---|---|---|---|
| S&P 500 | ^GSPC | index points | 3 Jan 2000 | 25 Sep 2026 | 6,723 |
| Amazon | AMZN | USD | 3 Jan 2000 | 25 Sep 2026 | 6,723 |
| Microsoft | MSFT | USD | 3 Jan 2000 | 25 Sep 2026 | 6,723 |
| Alphabet (Google) | GOOGL | USD | 19 Aug 2004 | 25 Sep 2026 | 5,561 |
| Oracle | ORCL | USD | 3 Jan 2000 | 25 Sep 2026 | 6,723 |
Method Stock prices are in US dollars and the S&P 500 in index points, all adjusted for splits and dividends. The first labelled day comes after the longest feature window has filled, and the last day is dropped because its outcome is not yet known.
Under the hood
How the site is built and kept current
Two layers
All data, feature, model, backtest and chart code lives in a plain Python library with no web framework in it, so the same functions run in a notebook, a test or this site. Django is a thin layer on top: each page validates its settings, calls one service function and renders. Charts reach your browser as Plotly JSON and are cached per ticker and training run.
Refreshed every weeknight
Nothing heavy runs while you browse. Each weekday at 22:30 UTC, after the US close, GitHub Actions runs the test suite, downloads the latest prices, retrains every model, builds the results into a container image and deploys it to Google Cloud Run. If a step fails, the site keeps the previous day's data. Latest training run: 28 Sep 2026, 00:33 BST.
Tested and typed
Unit tests run on synthetic prices with no network access. They include no-leakage tests, backtests checked against hand-computed returns, and regression tests for chart rendering quirks. The library is type-checked with mypy in strict mode and linted with ruff.
Built with
Python, pandas, NumPy, scikit-learn, statsmodels, Plotly, Django, WhiteNoise, yfinance and Parquet. The code is open: read it on GitHub.
Limits
Read the results with these in mind
Not investment advice
This is an educational project about method and communication. Nothing here recommends buying or selling anything.
Hindsight in the universe
These are today's giants, chosen today. Holding them since 2000 looks good partly because we already know they survived and grew, which flatters buy & hold.
Simple trading costs
A flat charge per position change stands in for spreads and commissions. There is no slippage, market impact, borrowing fee for shorts, or tax.
Trading at the close
A prediction made from a day's close is assumed to trade at that same close. In practice you would trade just before it or at the next open.
Video transcript
The video is a snapshot of the data to 23 Sep 2026. The live figures elsewhere on this page update with every nightly run.
- Title. The Up or Down logo, a green up-triangle over a red down-triangle, splits open at the gap that stands for today's close. "Will tomorrow's close be higher than today's? Predict it. Backtest it. Explain every step." A faint S&P 500 line runs up to today, then forks into "Up?" and "Down?".
- The data. "26 years of daily prices from 5 markets." Lines for the S&P 500, Amazon, Microsoft, Alphabet and Oracle draw from 2000 to 2026 on a log scale, each rebased to its first day, with market events marked. Alphabet ends at 136 times its 2004 listing price, Amazon at 56 times, Microsoft 14.2, Oracle 6.2 and the S&P 500 5.3 times their 2000 levels. "32,442 daily prices, adjusted for splits and dividends, cached as Parquet and refreshed every weeknight."
- Features. "Each day becomes 20 features, from the past only." A cursor moves through six months of the S&P 500 while everything to its right is hatched out as the hidden future. A card shows that day's RSI, MACD histogram, distance from the 30-day average, 21-day volatility, Bollinger position and volume ratio. On the last day, tomorrow is revealed: a fall, so the label is 0. "No look-ahead. Day t sees only data up to its own close, and tests change future prices to prove no earlier feature moves."
- Models. "7 models, scored only on days they never saw." Five expanding cross-validation folds fill in, then the final split: 80% of days to train and 20% held out as the test period, June 2021 to September 2026 (1,332 days). Each model's test ROC AUC on each market then appears as a dot around a dashed coin-flip line at 0.50. "Test ROC AUC runs from 0.48 to 0.53 across 35 model–market pairs. None of the models picked by cross-validation beat the naive baseline."
- Backtest. "Trade the predictions. Then pay the costs." The S&P 500's cross-validation pick, Random Forest, trades long/short against buy & hold over the test period. A cost slider moves: at 5 bp its Sharpe ratio is 0.64 against 0.69 for buy & hold, with free trading it is 0.80, and at 50 bp it falls to −0.74. "Long on predicted rises, short on predicted falls. The small edge disappears once trading costs are counted."
- Explore. "6 pages, and every chart is interactive." A browser scrolls through Overview, Indicators, Exploration, Models, Backtest and Multi-ticker, each listed with the question it answers, then switches to the dark theme as a phone layout slides in. "Light and dark themes, phone layouts too."
- Verdict. "Tomorrow's direction is close to a coin flip, and the site says so. What it offers instead: a leak-free method, honest baselines, and charts that explain every step." Up or Down, github.com/hashincludeim/ml-trading-backtester. Python, pandas, scikit-learn, Plotly and Django.