The same problem that exists for senators exists for presidents, only louder. There is no presidential WAR. There are mature, downloadable datasets and repeated expert surveys. What they measure is “how often did this person move the country in the direction the scorers already think is correct,” plus some contemporaneous popularity and some raw outcome numbers that a president only partly controls. “Was he any good?” remains a values question dressed up as a scoreboard.
What already exists
Historian and political-scientist rankings are the closest analogue to DW-NOMINATE plus interest-group scorecards. C-SPAN’s survey (ten equally weighted leadership categories, last run in 2021 after Trump’s first term) and Siena College’s (twenty categories, last full run 2022) are the two most cited. The top of both lists has been stable for decades: Lincoln, Washington, FDR, Theodore Roosevelt. Reagan sits ninth in the 2021 C-SPAN survey and lower (around 18th) in Siena 2022. Wilson has slid: he was routinely fourth-to-sixth in mid-20th-century polls and is now 13th in C-SPAN 2021. Trump’s first term landed 41st of 44 in C-SPAN 2021 and in the bottom three-to-five in Siena.
These surveys are not secret ballots of the American public. They are panels of academics and presidential specialists. The profession leans left; categories such as “pursued equal justice” and “moral authority” therefore move scores. Wilson’s drop tracks the heavier weight later generations placed on his resegregation of the federal civil service, the White House screening of The Birth of a Nation, and wartime civil-liberties statutes. Reagan’s relatively high C-SPAN placement reflects later credit for the 1980s recovery and the end of the Cold War; Siena’s lower placement reflects deficits, inequality, and Iran-Contra. Trump’s first-term scores reflect the same raters’ view of norms, rhetoric, two impeachments, and January 6 more than pre-COVID unemployment or the Abraham Accords. Public retrospective polls usually rank Reagan higher than the scholars do. That gap is data, not a rounding error.
Gallup (and successor) job-approval averages are the contemporaneous popularity series. Reagan averaged 52.8 percent over eight years, bottomed at 35 percent in the 1982 recession, and left office at 63 percent. Trump’s first term averaged 41 percent—the lowest of the polling era—and never reached 50 percent in Gallup. His second term, as of mid-2026, has also run well below 50 percent and remains the most polarized on record (Republican approval in the mid-80s, Democratic approval in single digits). Wilson predates regular Gallup; wartime support was high, then collapsed after his stroke and the Senate’s rejection of the League. Approval is not “goodness.” It is how the living public felt at the time.
Harder outcome series exist and are public: real GDP growth, unemployment and inflation paths, debt-to-GDP, stock-market returns, major legislation enacted, Article III judges confirmed, uses of force, treaties ratified. They do not self-interpret. Reagan inherited double-digit inflation and 7-plus percent unemployment, endured a sharp 1981–82 recession, then presided over a long expansion with inflation falling to the low single digits and unemployment to 5.5 percent; the publicly held debt roughly tripled. Trump’s first term saw the pre-COVID unemployment rate fall to 3.5 percent and a large corporate-rate cut; COVID produced a net job loss over the full term. Wilson’s first term produced the Federal Reserve, the income tax, the FTC, and the Clayton Act; the second produced wartime mobilization, the 1918–19 influenza interaction, and the 1920–21 depression. None of these numbers tell you whether the tax cut, the Fed, or the war was the right instrument.
Interest-group and think-tank scorecards (Heritage, ADA, etc.) exist for recent presidents the same way they do for senators. They grade against an explicit agenda. Useful once you have already chosen the agenda.
The three cases, measured rather than narrated
Wilson’s first-term legislative haul is large and durable: 16th Amendment, Federal Reserve, antitrust updates. His wartime leadership produced victory and then a peace settlement the Senate would not accept. His racial record and the Espionage and Sedition Acts are also durable. Later scholars lowered his rank as those last items received more weight; earlier scholars had treated the League vision as near-great even in failure. That is a change in the scoring rule, not new roll-call data.
Reagan’s measurable record is the disinflation, the 1980s expansion after the Volcker recession, the 1986 tax reform, three Supreme Court justices, and a defense buildup that coincided with Soviet overstretch and Gorbachev. The deficits, the rise in inequality, Iran-Contra, and the delayed federal AIDS response are also in the record. Historians still split on how much causal credit he deserves for 1989–91; the public has been more generous in hindsight.
Trump’s first term produced the 2017 tax law, three Supreme Court justices plus a large lower-court cohort, the Abraham Accords, territorial defeat of ISIS, no new major land war, and Operation Warp Speed. It also produced two impeachments (Senate acquittals), the COVID death toll and recession, the 2020 election challenge, and January 6. The second term is still being written; early 2025–26 data show heavy use of executive orders, tariff revenue, slower job growth than the late Biden years, and continued polarization. Historians have so far treated the first term as near-bottom; that judgment is only seven years old and was rendered by the same professional cohort that has been revising Wilson downward for different reasons.
What you would still need for a custom “performance” instrument
Exactly the same five choices required for senators, only the units change:
- A benchmark. Party median is useless for a president. You must pick: real per-capita consumption, peace, original-meaning originalism on the Court, “did the institutions he created still exist and work fifty years later,” predicted median-voter preference, or something else. The data will not choose.
- An action universe. Signed statutes only? Executive orders and national-security actions? Judicial appointments weighted by later doctrinal impact? Wars started versus wars ended?
- Weights. A Supreme Court seat or a world war is not a highway bill. Closeness of the congressional vote, contemporaneous public attention, and later scholarly citation are possible weights; each embeds a theory.
- Context. Inherited inflation, divided versus unified government, exogenous shocks (pandemic, oil shock, assassination of an archduke). Reagan’s 1982 recession and Trump’s 2020 collapse are not interchangeable with Wilson’s 1920–21 slump.
- Non-statutory activity. Rhetoric that shifts the Overton window, administrative capacity, scandal frequency, and “did later events make the decision look wise.” The last one is retrospective philosophy, not measurement.
Party-unity and first-dimension ideology scores work for the Senate because the denominator is clear: “voted with the team when the teams disagreed.” A president has no such clean denominator. That is why the historian surveys persist: they collapse the missing scoring rule into ten or twenty subjective categories and average the professors.
The honest limitation is identical to the one that opened the Senate discussion. Once you leave “how often did the economy grow / how many judges were confirmed / how many wars were avoided” and try to score “were they correct,” you are doing political philosophy or intergenerational judgment. The voting records, the economic time series, the appointment lists, and the approval numbers are excellent and public. The evaluation rule is the part that does not exist by consensus, which is why we will still be arguing about Wilson’s League, Reagan’s deficits, and Trump’s norms when the next set of historians is surveyed.

