Though the model performed well in OT25 it was not as strong as in the validation period. This is largely because the term skewed towards affirm relative to previous terms. Overall, the justices voted to reverse about 58 percent of the time; in the previous term, they voted to reverse more like two-thirds of the time. Case level outcomes similarly were relatively skewed to affirm relative to baselines.
This had two effects. The less important was a thresholding problem. The ideal probability threshold to separate reverse from affirm votes was not 0.5, but instead somewhat higher. The more important problem was that, general, the model has trouble detecting affirm votes. This was evident in the validation term, where the F1 score for the affirm class was markedly lower than for the reverse class. Last term, by leaning into affirm votes, the Court’s behavior went directly into the weaker part of the model’s ability, its relative blind spot. On a term with a normal affirm-reverse split, the model would likely have performed more like the validation term.
At the justice level, where did the model succeed and where did it fail? As shown in the figure, the model seems to have done best at the ideological extremes: its accuracies for Justices Sotomayor, Thomas, Gorsuch, and Jackson were on the order of 80 percent or higher. Justice Alito was just behind. By comparison, the model had the most trouble with those justices conventionally thought of as swing votes: Barrett, Kavanaugh, Roberts, and Kagan. Accuracies for those justices were on the order of 65 percent. Naturally, the cells become small once we look at the justice level, so the usual caveats apply.
justice level accuracy
So where does this leave us? Mainly, needing a better model of affirm votes. As luck would have it, I substantially revised the model for the upcoming term and believe I (largely) fixed the issue. More soon on that.