Up or Down About

Model comparison: Oracle

Trained on 5327 days (Apr 2000 to Jun 2021) and tested on the following 1332 days (Jun 2021 to Sep 2026). Every pipeline scales (and optionally applies PCA) inside each time-series fold.

Verdict

1 of 7 models beat the best naive guess

That guess is to always predict "up", which scores 51.7% test accuracy. The best ROC AUC is 0.526 (Bagging (KNN)); 0.5 means no better than chance. Daily direction is close to a coin flip, and these results reflect that honestly.

Test-set metrics

Will tomorrow's close be higher? CV figures are the mean ± std over expanding-window folds of the training period. Test figures come from the same fitted pipeline that is backtested.

ModelCV accuracyCV AUCTest accuracyPrecisionRecallF1ROC AUCvs naiveTuned params
Bagging (KNN) 49.3% ± 1.3% 0.496 51.7% 53.5% 50.0% 0.517 0.526 above defaults
Gradient Boosting 51.0% ± 1.4% 0.514 50.8% 52.5% 49.9% 0.512 0.504 below max_depth=2
Random Forest 51.5% ± 1.1% 0.526 48.6% 50.2% 49.1% 0.497 0.493 below max_depth=3
SVM (RBF) 51.5% ± 0.9% 0.516 48.5% 50.1% 52.2% 0.511 0.492 below C=0.3
Extra Trees 51.7% ± 0.7% 0.527 49.6% 51.1% 58.0% 0.543 0.489 below defaults
Linear SVC 51.5% ± 1.6% 0.504 49.3% 51.0% 49.4% 0.502 0.485 below defaults
Logistic Regression 50.7% ± 1.3% 0.516 48.9% 50.5% 53.6% 0.520 0.482 below C=0.01
Always up (baseline) –– 51.7%51.7%100.0%0.6810.500

Training period

How models scored during time-series cross-validation

Takeaway 4 of 7 models swing between beating and losing to a coin flip across folds, so their average CV score hides a lot of instability.

Test period

Unseen data: can the models rank up days above down days, and does any edge persist?

Takeaway Share of the test period each model spent above 50% (rolling 63 days): Logistic Regression 43%, SVM (RBF) 41%, Linear SVC 45%, Random Forest 43%, Extra Trees 50%, Bagging (KNN) 63%, Gradient Boosting 52%.

Walk-forward check

Would the result hold up if each model were refitted as new data arrived?

Takeaway Refitting 22 times changed test AUC by -0.003 on average (range -0.022 to +0.008). With 1,332 test days the 95% interval is about ±0.031 wide, so none of 7 models trained once and 1 walk-forward are reliably better than a coin flip. All 7 models are tested on the same days and largely agree, so one or two clearing the bar by luck is expected.

Method The models above are fitted once on the training period. Walk-forward keeps the same tuned settings but refits every 63 trading days (22 times) on all earlier days, then predicts only the next block. The backtest still uses the model trained once.

Inside Bagging (KNN)

Pick another model from the selector or the table above

Takeaway Bagging (KNN) has a test ROC AUC of 0.526: pick a random up day and a random down day, and it ranks the up day higher 53% of the time (50% = guessing).

Method Permutation importance is measured on the last validation fold of the training period, never on the test set. It's always reported for the original named features, even for PCA pipelines.