Psychology
Psychology

Levels of Evolutionary Analysis and Hypothesis Generation

Psychology

Levels of Evolutionary Analysis and Hypothesis Generation

You're convinced that jealousy is an evolved adaptation, not a culturally constructed Western neurosis.
developing·concept·1 source··May 10, 2026

Levels of Evolutionary Analysis and Hypothesis Generation

You Want to Test Whether Jealousy Is Adaptive. Where Do You Start?

You're convinced that jealousy is an evolved adaptation, not a culturally constructed Western neurosis. You want to test it. What do you do?

You can't test "jealousy is evolved" directly. The claim is too broad. Selection-by-natural-selection is not a hypothesis you test in a laboratory; it's the framework everything else sits inside. Working downward from there, you need a more specific hypothesis. Trivers's parental investment theory predicts the sex that invests more in offspring will be choosier about mates and the lower-investing sex will compete more for sexual access. That's tighter, and it's testable, but it's still a theory rather than a specific hypothesis about jealousy. Working downward again: Men face paternity uncertainty that women do not, so men's jealousy should be especially activated by sexual rather than emotional infidelity. That's a specific evolutionary psychological hypothesis. Working downward one more time: In a forced-choice experiment, men should report sexual infidelity as more distressing than emotional infidelity at higher rates than women report the same. That's a prediction. You can run it. You can falsify it.

What just happened was four moves through Tooby and Cosmides's hierarchy of evolutionary analysis: from general evolutionary theory through middle-level theory through specific hypothesis to specific prediction.1 Buss treats this hierarchy as the discipline's core methodological frame. It is how evolutionary psychology generates testable claims out of an evolutionary metatheory that is itself not directly testable in any single study.

The hierarchy is one of two strategies. The other runs the opposite direction — start from an observation and reverse-engineer back to evolutionary function. Singh observed that men across cultures rate women with low waist-to-hip ratios as more attractive. Working backward, he hypothesized that the preference is an adaptation for detecting fertility cues, since low WHRs in fact correlate with health and fertility outcomes. He then ran specific predictions to test the hypothesis.2 Both strategies — top-down theory-driven and bottom-up observation-driven — produce testable claims when done carefully. Buss treats them as complementary rather than competing.

Definition / Core

Pin down the four levels.

General evolutionary theory sits at the top. The framework includes natural selection, sexual selection, inclusive fitness theory, the modern synthesis combining genetics with selection, and the broader claim that complex functional design in organisms is produced by selection over many generations. Most working researchers treat general evolutionary theory as effectively true and use it as a constraint on lower-level hypotheses. Buss notes that the framework is in principle falsifiable — by, for example, the discovery of complex life forms produced too rapidly for selection to have built them, or adaptations functioning solely for the benefit of other species, or adaptations functioning for the benefit of same-sex competitors — but no such observations have ever been documented.3 In practice, general evolutionary theory functions as background.

Middle-level evolutionary theories sit one level down. These are still broad — they cover entire domains of biological functioning — but they make specific claims that can be tested and potentially falsified.4 Trivers's parental investment theory is the textbook example. It claims that the sex that invests more in offspring will evolve to be choosier in mate selection, and the sex that invests less will compete more for sexual access to the high-investing sex. The theory is not derivable from general evolutionary theory; nothing in natural selection by itself implies parental investment. The theory has to stand or fall on its own merits, supported by empirical data from many species. The cumulative weight of evidence — including the so-called sex-role-reversed species like the Mormon cricket and the pipefish seahorse, where males invest more and females compete more — supports parental investment theory broadly.4 Other middle-level theories include kin-selection theory (Hamilton's rule), reciprocal altruism theory (Trivers 1971), and sexual selection theory (Darwin 1871).

Specific evolutionary hypotheses sit one level further down. These propose specific psychological mechanisms in specific organisms with specific design features. The hypothesis that women have evolved specific preferences for men with resources is specific.5 It can be tested. Predictions follow: women should value qualities linked to resource acquisition (status, intelligence, ambition); women's attention in dating contexts should be drawn more to resource-cued men; women whose partners fail to provide resources should divorce them at higher rates. Each prediction is testable, and the cumulative weight of the predictions' empirical fates determines whether the hypothesis survives.

Specific predictions sit at the bottom. These are the empirically testable claims that follow from the specific hypothesis. Predictions can fail without invalidating the hypothesis (perhaps the relevant condition was missing in the test population). The hypothesis can be wrong even if the middle-level theory is right (perhaps the relevant mutations didn't arise in the right population). And the middle-level theory can be wrong even if general evolutionary theory is right (perhaps Trivers got parental investment wrong but selection still works as Darwin described). The hierarchy means that falsifying a low-level prediction does not automatically falsify the higher levels above it.6

The two hypothesis-generation strategies:

Top-down (theory-driven) moves from general theory to specific predictions. Start with parental investment theory, derive the hypothesis that women will be choosier in long-term mate selection, predict that women will impose a longer delay before consenting to sex than men, test the prediction empirically. Buss and Schmitt 1993 documented exactly this — women impose longer delays and more stringent standards before sexual consent than men.7 The strategy works when the theory is mature enough to make specific predictions that wouldn't otherwise be obvious.

Bottom-up (observation-driven) moves from observation to functional hypothesis. Observe that men prioritize physical appearance more than women in mate selection, hypothesize that physical appearance provides cues to fertility that men's preferences evolved to track, derive predictions about which specific physical features should be most attractive (low WHR, facial symmetry, youth markers), test the predictions. Singh's WHR research is the textbook example.8 The strategy works when the observation is robust but the underlying function is not obvious from theory alone. Buss frames it as reverse engineering — running design analysis backward from observed phenomena to inferred function.

Buss positions both strategies as part of standard scientific practice. Astronomy used a similar pattern when the expanding universe was observed first and theories to explain it followed. The bottom-up strategy is not less rigorous than the top-down one; it just works in a different direction.9

Evidence

The evidence for the hierarchy is the discipline's productivity over its first three decades. Multiple specific hypotheses generated from middle-level theories have produced robust cross-cultural empirical findings. Buss's own 37-cultures study (10,047 participants) tested specific predictions from parental investment theory across radically different cultural contexts.10 Women across all 37 cultures preferred mates with greater financial resources, higher social status, and greater age — predictions that follow from the middle-level theory and that hold under test.

Sexual jealousy research provides another worked case.11 The hypothesis: men's jealousy should be especially activated by sexual infidelity (paternity uncertainty), women's jealousy by emotional infidelity (resource diversion). The prediction: in forced-choice experiments asking which would be more distressing, men should select sexual infidelity at higher rates than women. The result: 60 percent of men vs 17 percent of women select sexual infidelity as more distressing; 83 percent of women vs 40 percent of men select emotional infidelity. The pattern replicates across the United States, Korea, Japan, the Netherlands, Sweden, Brazil, England, Romania, and dozens more cultures. fMRI evidence shows different brain activation patterns by sex during imagined infidelity.11 The hypothesis survives because the predictions hold up.

Methods tested at the bottom of the hierarchy span eight categories that Buss tabulates in his textbook (Table 3): comparing different species, cross-cultural methods, physiological and brain imaging methods, genetic methods, comparing males and females, comparing individuals within a species, comparing the same individuals in different contexts, and experimental methods.12 Six categories of data sources stack with these: archeological records, hunter-gatherer studies, observations, self-reports, life-history data, and human products. The methodology is not designed for any single decisive test; it is designed for convergent evidence across multiple methods that don't share their methodological limitations. A finding that holds in self-report and replicates in physiological measurement and shows up in cross-cultural data and matches what archeology suggests is a finding with multiple independent supports.

Identifying which adaptive problems exist in the first place — the input to the whole hierarchy — uses six guidelines.13 Modern evolutionary theory's broad classes (survival, mating, parenting, kin investment) supply a first cut. Universal human structures (group living, status hierarchies) supply a second. Hunter-gatherer studies, which approximate ancestral conditions, supply a third. Paleoarcheology, where bones and stones leave evidence of ancestral diet, injury, and tool use, supplies a fourth. Current psychological mechanisms, where existing phobias and preferences provide windows into ancestral hazards, supply a fifth. Task analysis (Marr 1982), where the cognitive and behavioral tasks required for an observed phenomenon are decomposed, supplies a sixth.14 No single guideline is sufficient; together they triangulate.

Marr's contribution deserves separate attention. Task analysis decomposes a phenomenon into the steps it requires and asks what cognitive mechanisms each step needs. To explain why people aid genetic relatives more than non-relatives, you need a kin-recognition mechanism (how do they know who's related?) and a kinship-closeness mechanism (how do they tell close kin from distant?). Each subtask suggests the design features of the underlying adaptation. Marr's framework was originally developed for vision, but the analytical move generalizes — and Buss explicitly imports it as a tool for identifying psychological adaptations.14

Tensions

The hardest tension in the hierarchy is the relationship between falsification at different levels. A specific prediction can fail without falsifying the specific hypothesis (perhaps the conditions in the test population were unusual). The specific hypothesis can fail without falsifying the middle-level theory (perhaps the relevant mutation didn't arise in this lineage). The middle-level theory can fail without falsifying general evolutionary theory (perhaps Trivers got parental investment wrong, but natural selection still works). Buss treats this not as a bug but as a feature — the hierarchy lets researchers preserve productive frameworks even when individual predictions fail.6

Critics have argued the feature is also a bug. If failed predictions never falsify higher levels, the discipline can absorb arbitrary amounts of disconfirming evidence without updating. The response from Tooby, Cosmides, and Buss is that cumulative failure of predictions does count against higher levels. Sustained, broad-based failure of predictions from a middle-level theory eventually falsifies the theory. The mate-deprivation hypothesis of male sexual coercion is one example of a hypothesis that didn't survive — failed predictions across multiple studies eventually retired it.15 The hierarchy doesn't prevent falsification; it just requires the failure to be cumulative.

A second tension runs around bottom-up reverse engineering. The bottom-up strategy is more vulnerable to just-so storytelling. Once you observe a phenomenon, you can usually construct some evolutionary story about why selection would have built it. The story might be right or wrong. Distinguishing legitimate reverse engineering from post-hoc rationalization requires the researcher to derive predictions from the proposed function and test those predictions. Singh's WHR work passed this test — the hypothesis that low WHR cues fertility predicted specific cross-cultural patterns that were then confirmed empirically.8 Other cases haven't passed it as cleanly. The bottom-up strategy is not bad methodology; it's methodology that requires more discipline to do well.

A third tension is about middle-level theories themselves. Some are well-supported and broadly accepted (kin selection, parental investment, reciprocal altruism). Others are contested (group selection, multilevel selection theory). Working researchers gravitate toward the well-supported theories, which means specific hypotheses generated from contested middle-level theories get less attention. The discipline's empirical agenda is partly shaped by which middle-level theories are currently in vogue, which can produce blind spots.

A fourth tension is about the boundary between general evolutionary theory and middle-level theories. Inclusive fitness theory is sometimes treated as middle-level (it's specific enough to make testable predictions about kin-directed altruism) and sometimes as general (it's part of the basic framework that selection operates). Where the line gets drawn affects how researchers frame their work, and the line is not always clear.

A fifth tension is methodological. The convergent-evidence principle works only when the methods being convergent don't share methodological limitations. Self-report data, observer-report data, and behavioral data sometimes converge because they share a common bias, not because they're independently confirming the same finding. Distinguishing genuine convergence from shared-method convergence requires care.

Author Tensions & Convergences

Tooby and Cosmides supplied the formal framing of the hierarchy. Their 1992 chapter in The Adapted Mind established the levels and the falsification logic Buss adopts. The convergence between Buss and Tooby/Cosmides on the methodological hierarchy is total. Where Buss adds emphasis is on its application — most of his textbook is structured around running specific hypotheses through the hierarchy, demonstrating the methodology rather than just describing it.

Trivers operates at the level of middle-level theories. His three seminal contributions — reciprocal altruism (1971), parental investment theory (1972), and parent-offspring conflict (1974) — are the textbook examples of middle-level theories that productively generate specific hypotheses across many domains.16 Buss treats Trivers as the model contributor at this level. The convergence shows up in how Buss organizes mating, parenting, and cooperation chapters around Trivers's theories. The disagreement, where it exists, is about specific applications — for example, exactly how parent-offspring conflict applies to human weaning or to mother-fetus interactions.

Marr supplied task analysis as a methodological tool, originally for vision, that Buss imports for evolutionary psychology more broadly.14 The convergence is on the analytical move: decompose an observed phenomenon into the cognitive and behavioral tasks it requires, then identify the design features of the mechanisms each task needs. Marr's three levels of analysis (computational, algorithmic, implementational) and Tinbergen's four whys (immediate, developmental, functional, phylogenetic) overlap but are not identical. Buss uses Marr's task analysis at the design-features level without integrating Marr's broader three-levels framework.

Where Buss diverges from his predecessors is mostly in tone. Tooby and Cosmides write polemically; the strongest version of the EP case is the version they make. Buss writes pedagogically; he reports the strong version, the weak version, and the empirical state of the dispute, and lets the reader weigh. The textbook is a work of synthesis rather than advocacy. Some readers will find this judiciousness valuable. Others will find it occasionally too even-handed about disputes where the data favor one side decisively.

A separate tension runs to the broader scientific methodology literature. Lakatos's framework of research programs with hard cores and protective belts maps onto Buss's hierarchy in suggestive ways. General evolutionary theory is the hard core; middle-level theories are part of the protective belt; specific hypotheses sit at the periphery. Lakatos's framework explains why scientific frameworks survive individual disconfirmations: the protective belt absorbs the damage. Buss does not name Lakatos, but the structural fit is close. The vault implication: evolutionary psychology is methodologically a Lakatosian research program, and its productivity should be evaluated by whether it generates novel predictions that get confirmed (which it has) rather than by whether individual predictions ever fail (which they do).

Cross-Domain Handshakes

The hierarchy of analysis maps onto Marr's framework for AI and cognitive science in a way that produces a working pedagogical bridge. Marr's three levels — computational, algorithmic, implementational — were developed for vision but generalize to any computational system. The computational level asks what problem is being solved and why. The algorithmic level asks what procedure is being run to solve it. The implementational level asks what physical substrate runs the procedure.14

Compare to Buss's hierarchy. Middle-level theories operate at Marr's computational level — they specify the adaptive problems being solved. Specific evolutionary hypotheses operate at Marr's algorithmic level — they specify the procedures (decision rules, design features) the mechanism uses. Predictions and methods operate near Marr's implementational level — they specify the observable outputs the mechanism produces in particular substrates and contexts. The frameworks are not identical, but the structural correspondence is real.

For AI-assisted creator work and for general thinking about AI systems, the implication is that any complete account of an AI system needs work at all three Marr levels — what is the system trying to do, what algorithm is it running, and what substrate runs the algorithm. This maps onto evolutionary psychology's hierarchy in a way that gives both fields a shared methodology. The insight neither domain generates alone: cognitive analysis at multiple levels is a general feature of explaining computational systems, whether evolved or trained, and the frameworks built independently in EP and in AI converge on similar structural commitments because they're addressing the same underlying problem of how to explain systems with internal complexity.

A second handshake runs to behavioral-mechanics, where understanding the levels of analysis equips an operator to choose the right intervention point. Most BM techniques are prediction-level interventions — present a specific cue, get a specific response. The technique catalog is organized around inputs and outputs without explicit reference to the underlying mechanism. The operator who knows the levels of analysis can design new interventions by working back through them. What adaptive problem does the target's mechanism solve? (Middle-level theory.) What design features must the mechanism have? (Specific hypothesis.) What inputs would activate those design features in this specific situation? (Prediction-level intervention.)

The connection to existing BM pages: the manipulation and influence hub catalogs techniques without explicit middle-level-theory framing. The same hub re-read with the levels of analysis becomes a more generative resource. Operators who understand parental-investment theory know which kin-related mechanisms can be triggered. Operators who understand reciprocal-altruism theory know which cooperation-related mechanisms can be triggered. The catalog is converted from a set of techniques to be memorized into a set of predictions derivable from middle-level theories. The insight neither domain generates alone: technique catalogs become principled rather than memorized when annotated with which middle-level theory each technique exploits, and new techniques can be derived from middle-level theories rather than discovered by trial and error.

The Live Edge

The Sharpest Implication.

When you encounter an evolutionary psychology claim — in popular journalism, in self-help, in political discourse — most of the work of evaluating it happens at the right level. A bad claim is often a claim made at the wrong level for the kind of evidence presented.

If someone says "evolutionary psychology proves that men are naturally promiscuous," they're collapsing the levels. The middle-level theory of parental investment predicts sex differences in mating strategy on average; it does not predict that any specific man will be promiscuous, and it does not say anything about what any individual man should do. Treating a population-level prediction as a personal-trait claim is a category error. The prediction operates at a different level than the claim about individual behavior.

If someone says "evolutionary psychology is just storytelling because the predictions can't be falsified," they're missing the hierarchy. Specific predictions are routinely falsified — the mate-deprivation hypothesis of sexual coercion was falsified by data and retired. What survives the cumulative weight of evidence is the higher-level theory the predictions came from. The claim that the framework can't be falsified is correct only at the very top level (general evolutionary theory), where falsification would require things that have never been observed. At every level below that, falsification is routine.

Knowing the hierarchy lets you do something more productive than agree or disagree with EP claims. You can locate which level the claim is at, ask what predictions follow from it, and check whether the predictions hold. Most poorly framed EP claims dissolve when you ask "what specific prediction follows from this, and what would falsify it?" Most well-framed claims survive that question and hold up against evidence.

Generative Questions.

If middle-level theories function as the protective belt of an evolutionary research program, what would a comprehensive map of which middle-level theories are well-supported, contested, and falsified look like? The discipline does not maintain such a map publicly. Vault work could.

The bottom-up strategy is more powerful for novel domains where theory is underdeveloped. The top-down strategy is more powerful for mature domains where theory is well-specified. Where in evolutionary psychology should each strategy currently dominate, given the state of theory in different sub-domains?

Marr's three levels and Tinbergen's four whys and Buss's four levels of analysis all carve up the explanatory landscape differently. Are they three different frameworks for the same territory, or three different territories that happen to overlap? The answer affects how the vault should connect them.

Connected Concepts

Open Questions

  • How much disconfirming evidence at the prediction level is required before a middle-level theory is updated or retired? The discipline has no formal threshold; informal practice varies.
  • The bottom-up strategy is essential for novel domains but vulnerable to just-so storytelling. What discipline-level checks distinguish productive reverse engineering from post-hoc rationalization?
  • The convergent-evidence principle assumes methods don't share limitations. When self-report, observation, and physiology all converge, are they confirming the same finding or sharing a common bias? Distinguishing requires meta-methodological work that the discipline does occasionally but not systematically.
  • Marr's three levels and the EP hierarchy converge on similar structures. Is this a deep correspondence or a superficial one? Mapping it carefully could yield a unified meta-framework for explaining evolved and trained computational systems.

Footnotes

domainPsychology
developing
sources1
complexity
createdMay 10, 2026
inbound links4