Model comparison: Alphabet (Google)
Trained on 4397 days (Nov 2004 to May 2022) and tested on the following 1100 days (May 2022 to Sep 2026). Every pipeline scales (and optionally applies PCA) inside each time-series fold.
Verdict
0 of 7 models beat the best naive guess
That guess is to always predict "up", which scores 52.8% test accuracy. The best ROC AUC is 0.522 (Logistic Regression); 0.5 means no better than chance. Daily direction is close to a coin flip, and these results reflect that honestly.
Test-set metrics
Will tomorrow's close be higher? CV figures are the mean ± std over expanding-window folds of the training period. Test figures come from the same fitted pipeline that is backtested.
| Model | CV accuracy | CV AUC | Test accuracy | Precision | Recall | F1 | ROC AUC | vs naive | Tuned params |
|---|---|---|---|---|---|---|---|---|---|
| Logistic Regression | 51.0% ± 2.4% | 0.498 | 52.5% | 53.5% | 77.1% | 0.631 | 0.522 | below | C=1.0 |
| Linear SVC | 51.0% ± 2.4% | 0.497 | 52.5% | 53.5% | 77.8% | 0.634 | 0.522 | below | defaults |
| Gradient Boosting | 51.0% ± 1.2% | 0.506 | 50.8% | 53.2% | 57.5% | 0.553 | 0.518 | below | max_depth=5 |
| SVM (RBF) | 50.5% ± 2.3% | 0.507 | 51.8% | 53.2% | 73.7% | 0.618 | 0.516 | below | C=1.0 |
| Extra Trees | 51.0% ± 2.5% | 0.495 | 52.2% | 52.6% | 96.2% | 0.680 | 0.508 | below | defaults |
| Random Forest | 51.0% ± 1.1% | 0.499 | 51.1% | 52.3% | 84.7% | 0.647 | 0.507 | below | max_depth=6 |
| Bagging (KNN) | 51.9% ± 2.9% | 0.515 | 50.9% | 53.0% | 62.5% | 0.573 | 0.499 | below | defaults |
| Always up (baseline) | – | – | 52.8% | 52.8% | 100.0% | 0.691 | 0.500 |
Training period
How models scored during time-series cross-validation
Takeaway 7 of 7 models swing between beating and losing to a coin flip across folds, so their average CV score hides a lot of instability.
Test period
Unseen data: can the models rank up days above down days, and does any edge persist?
Takeaway Share of the test period each model spent above 50% (rolling 63 days): Logistic Regression 68%, SVM (RBF) 71%, Linear SVC 68%, Random Forest 53%, Extra Trees 65%, Bagging (KNN) 54%, Gradient Boosting 53%.
Walk-forward check
Would the result hold up if each model were refitted as new data arrived?
Takeaway Refitting 18 times changed test AUC by +0.010 on average (range -0.010 to +0.023). With 1,100 test days the 95% interval is about ±0.034 wide, so none of 7 models trained once and none walk-forward are reliably better than a coin flip.
Method The models above are fitted once on the training period. Walk-forward keeps the same tuned settings but refits every 63 trading days (18 times) on all earlier days, then predicts only the next block. The backtest still uses the model trained once.
Inside Bagging (KNN)
Pick another model from the selector or the table above
Takeaway Bagging (KNN) has a test ROC AUC of 0.499: pick a random up day and a random down day, and it ranks the up day higher 50% of the time (50% = guessing).
Method Permutation importance is measured on the last validation fold of the training period, never on the test set. It's always reported for the original named features, even for PCA pipelines.