Most panels mistake confidence for competence. The Army is no exception.
Picture a selection panel: five experienced leaders around a table, choosing the next leader of a mission-critical organization. The candidate walks in and owns the room inside 90 seconds. He is polished and quick. He tells a terrific story about a crisis he turned around, decisive and calm at the center of his own narrative. Asked about a time things went badly, he reframes the failure into a triumph in about eight seconds, and the panel chuckles. The most senior member leans back, nods and says, “Now that’s a leader.” Consensus is arrived at in 40 minutes.
There was another candidate, and the panel had half-forgotten him already. He asked clarifying questions before he answered. He paused to think. At one point he said, plainly, “I got that one wrong, and here is what I learned.” Next to the first candidate, he’s read as tentative. Less of a presence.
The first candidate got the job. And the panel never gathered a shred of structured evidence about the only question that mattered: which candidate was right for this organization, this mission, and this moment? That scene is not a caricature. It is, more or less, how the army, and most American institutions, select many of their most consequential leaders.
The process feels rigorous. Files are read, records compared, hard questions asked. But strip away the ceremony and what remains is a room rewarding the performance of leadership rather than evidence of it. Most panels mistake confidence for competence. The army is no exception.
“Best” Is a Ranking. “Right” Is a Match.
When one calls an officer the best, one almost always means a ranking in the abstract: the strongest file, the most impressive record, the top of the order of merit list. That is a property of a person, measured independently of any particular job. Right is a different kind of claim. It is a match between a person and a context…this mission, this formation, this command climate, this moment in the fight. The best leader optimizes a person. The right leader optimizes a fit. Put a brilliant, decisive crisis commander in charge of a fragile coalition negotiation and watch him struggle; not because he got worse, but because the context changed and the match broke.
Why do selection processes get this wrong so reliably? Because of one of the most durable findings in leadership research: the traits that cause someone to emerge as a leader are not the traits that make them effective in the role. Confidence, dominance, verbal fluency, and charisma predict who rises, who wins the room, and who interviews well. Narcissism, for instance, reliably predicts who gets seen as a leader and has no relationship at all with how well they will lead. Effectiveness runs through different terrain: judgment, adaptability, humility, and fit to the job’s real demands. An unstructured board is, functionally, an emergence detector. It rewards the officer who looks like a leader and systematically under-selects the one who would actually deliver. The board chooses the best-looking candidate, not the right one.
The Science Just Moved
For a generation, the map of selection science seemed settled. In 1998, Frank Schmidt and John Hunter synthesized 85 years of research and ranked selection methods by how well they predict job performance. General cognitive ability sat near the top, with a validity around .51. For readers who do not live in statistics: validity runs from zero, a method that predicts nothing, to 1.0, a method that predicts perfectly, and in personnel selection anything above roughly .35 is considered very useful. A .51 was not merely useful; it was miles ahead of most alternatives, with unstructured interviews near .38 and years of education near .10. The practical lesson was simple and seductive: if you can measure only one thing, measure raw intellect. The smartest candidate is the best bet.
In 2022, that map was redrawn. Paul Sackett and colleagues demonstrated that the field’s canonical estimates had systematically overcorrected for a statistical artifact called range restriction, inflating validity across the board. When they recomputed, two things happened. The numbers came down and the rankings reordered. Cognitive ability fell from roughly .51 to .31. The structured interview rose to the top, at roughly .42. Read that carefully: the field’s favorite proxy for best (raw intellectual horsepower) is no longer the standout predictor. The strongest single method is the one that forces an organization to define what a specific job actually demands and then score observed behavior against those demands. The evidence itself now favors right, which is anchored in context, over best, which sails above it.
There is a second, older lesson buried here. Since Paul Meehl’s work in the 1950s, confirmed by meta-analysis ever since, scholars have known that structured combinations of valid data outperform expert intuition at prediction. Experts are excellent at diagnosis and unreliable at prognosis. “I know a leader when I see one” is the least reliable instrument in the room, and in an unstructured board, it is often the only one in use.
A Case Study: The Command Assessment Program
The U.S. Army has already tested this science on itself, and the case deserves to be read for exactly what it is: one example of a problem that runs through every selection board the service convenes. Beginning as a pilot in 2019 and running at full scale from 2020-2025, the Command Assessment Program put battalion and brigade command candidates through exactly the discipline the research prescribes: multiple validated instruments, structured interviews, behavioral observation by trained assessors, peer and subordinate feedback, and decision rules set before anyone met a candidate. Nearly 2,000 candidates were assessed a year. The army discontinued the program in 2025 and returned command selection to traditional boards.
A selection program founded on the premise of evidence must produce evidence about itself. Without it, no design pedigree will save the program, and, by its own standard, none should.
Two lessons survive the program. The first is that procedural discipline is feasible: an institution the size of the Army built a context-anchored, multi-method assessment and ran it at scale for six years. Special operations forces have selected leaders this way for decades and still do, watching how candidates behave under stress, ambiguity, and fatigue rather than trusting the file to speak. The second lesson is harder, and selection professionals should take it squarely: in six years, the Command Assessment Program never published the validation evidence that would have demonstrated its value; no reported link between assessment results and later performance in command. When scrutiny came, it had a sound design but no empirical defense. That is on the professionals who build such systems. A selection program founded on the premise of evidence must produce evidence about itself. Without it, no design pedigree will save the program, and, by its own standard, none should.
The Problem Scales
Strip away the program’s name and what remains is the Army’s default: a board of experienced people, a file, a short conversation, and an impression. That default governs far more than battalion command. Non-commissioned officer (NCO) promotion and selection boards, by the account of nearly anyone who has sat on one, reward bearing, polish, and the ability to answer under pressure, all cousins of the confidence vignette described earlier. In the reserve components, where boards are local, personalities are known, and the pool is thin, selections vary so widely from unit to unit that factors other than merit routinely decide them. The error is the same at every board; only the cost changes.
And the cost peaks precisely where the process never reached. Many general and flag officers were assessed earlier in their careers, at selection for command or for specialized units. But the selection of generals and admirals itself, the boards and slating decisions that put a name behind a three- or four-star desk, has never included psychological assessment or structured interviews at all. The army screened on the way up and never screened where the stakes are highest.
The stakes at that level are not abstractions. A senior leader shapes command climate, trust, retention, and the decision quality of everyone below them, and then chooses who gets developed next. When senior leaders fail, they rarely fail on technical grounds. The failures are overwhelmingly interpersonal and ethical: precisely the domains that files and unstructured interviews measure worst and that validated behavioral assessments measure best. The same three questions apply to the NCO and the general alike: Can they do the job? Will they do the job? And do they fit this context?
Measure Merit, Then Defend It
The path forward begins by taking the current standard seriously. The U.S. Army has been clear that selection and promotion should rest on merit and performance. I agree, and that is an argument for measurement, not against it. Merit is precisely what structured, validated assessment captures best. Reputation, presence, and unstructured board impressions are where favoritism and bias have always lived.
The institutional resistance to that idea is real, and any recommendation that ignores it is a wish. So, this one is built to survive it. First, start where the stakes are lowest and the volume is highest: structured, behaviorally anchored interview protocols cost almost nothing to add to NCO boards, and the thousands of selections made each year would generate validation data faster than any officer program could. Second, make structure the default rather than the innovation: training in anchored scoring should be a standard part of board member preparation. Third, write the proof requirement into the charter: any assessment effort should be required from its first day to define what success in the job looks like, track selected leaders against it, and publish the result on a fixed schedule, so that it can defend itself when scrutiny comes, as it will! And fourth, keep the decision with leaders who own it: assessment should arrive as decision support for a board, never as a gate that replaces it, which removes the most common objection before it is raised.
The army is not short on science. The strongest evidence in a century of selection research points this way. It is not short on proof of concept. The army ran the process at scale for six years, and its special operations forces still do. What it has been short on is the habit of proving what works, and the institutional patience to apply its own best practice to its most consequential choices. The army does not need a better way to find its best leaders. It needs the discipline to find the right ones…and the evidence to show it.
Clayton Manning is an Assistant Professor in the Department of Command Leadership and Management at the United States Army War College and an operational psychologist with over 20 years of experience advising senior leaders in complex, high-consequence environments. He has designed, led, and evaluated leader assessment and selection programs within some of the military’s most demanding organizations, while also serving in executive leadership roles, including as a hospital commander and chief of operations for a directorate of psychological applications.
The views expressed in this article are those of the author and do not necessarily reflect those of the U.S. Army War College, the U.S. Army, or the Department of War.
Photo Description: Candidates from cohort 5 attempt to traverse an obstacle at the Leader Reaction Course during the Battalion Commander Assessment Program January 23, 2020, at Fort Knox, Ky. More than 800 officers will complete cognitive and non-cognitive, physical, verbal and written assessments that will provide a more holistic look of an officer before being selected for battalion command.
Photo Credit: U.S. Army photo by Staff Sgt. Daniel Schroeder, Army Talent Management Task Force