Pick an essayist you've read enough of. Joan Didion. James Baldwin. Annie Dillard. Open one of their essays at random. Read three sentences. You can feel it's them. You couldn't necessarily say why — but you could pass a blind test.
What is that recognition? The standard answer is voice — and the standard story about voice is romantic. Voice is a literary thumbprint waiting to be discovered. Each writer has one and only one. Voice is ineffable, magical, mostly genetic, requires years of patient excavation to find.
Michael Deen rejects every word of this. His position, laid out in the Voice essay that umbrellas the Spirit/Sound/Sight half of his framework:
"I don't even like the romantic idea that each writer has their own precious voice, a literary thumbprint waiting to be discovered, that unifies all their work with an iconic aura. Your voice should be constantly evolving with you, not frozen into a brand."1
Voice is not a thumbprint. Voice is the current configuration of nine engineerable axes. Change the configuration, change the voice. There is no inner essential voice underneath that you'd find if you scraped enough to reveal it. The configuration is the voice.
This is the move that ties the whole Voice half of Essay Architecture together.
If Spirit (7) + Sound (8) + Sight (9) compose the Voice half, then Voice = the simultaneous output of nine axes:
From Spirit: Tone, Perspective, Subtext From Sound: Repetition, Rhythm, Rhyme From Sight: Imagery, Words, Motif
Each axis has a current setting. Each setting can move. The reader experiences the combined output and recognizes it as a particular writer's voice.
The recognition test (you-feel-it's-them-in-three-sentences) works because in any three-sentence sample, the writer's settings on most of the nine axes will be visible. Tone register. Sentence rhythm. Word choice. Image density. Subtext depth. The combination is rarely repeated by another writer, so the combination functions as identity even though it isn't essential identity.
The thumbprint metaphor implies stability — a fingerprint doesn't change. Deen's claim is that voice should change. "Your voice should be constantly evolving with you, not frozen into a brand."1
A writer who treats voice as brand will optimize for consistency — produce more of the same recognizable thing, build a recognizable signature, get marketed as that signature. The reader trains to expect the signature. The writer trains to deliver it. The signature ossifies. The writer becomes a brand asset.
The cost of brand-voice: when life changes the writer (illness, divorce, conversion, age), the writing can't follow without seeming inconsistent. The reader who came for the brand feels betrayed by evolution. The writer who refuses to evolve produces work increasingly disconnected from their actual interior.
Deen's alternative: treat voice as the live-time settings you're choosing this paragraph. Yesterday's settings were for yesterday's material. Today's material may need different settings.
Standard voice-criticism uses adjectives. "Casual." "Funny." "Abrasive." "Self-conscious." Deen rejects these as too reductive. "It is too reductive to define voice with petty adjectives."1
An adjective collapses nine axes into one label. "Casual" might mean: low register + low confidence + active voice + 2nd-person inflection + low-density imagery + short rhythm + minimal subtext. That's seven different settings collapsed into one word.
The result of adjective-criticism: writers try to be casual or funny or whatever the label is, and they over-pull on whichever axis they think the label most maps to. Comedians who decide they have a "dry" voice over-pull on understatement and lose tonal range. Memoirists who think their voice is "vulnerable" over-pull on emotional disclosure and lose argumentative reach.
Deen's alternative: drop the adjective entirely. Think about the nine axes. Each one is doing something specific. Each one can be tuned independently.
The reason Voice matters at all: reading is hard. "We can't forget that reading is an act of friction: it takes work to compile strings of words into meaning, and so every sentence is cognitive labor. To combat this, you want to hit your readers in their nervous system."1
Voice is the layer that makes reading less effortful. "Just as important as what you say (logos) is how you say it (pathos)."1 Voice carries the pathos. It compels.
The canonical authorities Deen invokes: Susan Sontag and Joseph Conrad both said the purpose of voice is to activate the senses — make you feel, make you hear, make you see. "Good prose turns language into a linguistic hallucination; it compels a reader to read thousands of words on any topic."1
Linguistic hallucination. This is the target. The reader is no longer parsing — they're seeing, hearing, feeling along with the prose. Voice is what transforms the reading experience from cognitive labor into something closer to sensory perception.
Deen frames the AI question directly. "As the Internet gets crowded and AI advances, less writers will doubt the importance of voice; the real debate is on where voice comes from. Is it supernatural or can it be captured in a style guide? Is it artistic or engineered? Spontaneous or imitated? Does editing kill it or augment it? I don't love these questions."1
The framing matters: the importance of voice is now uncontested because AI writes prose without voice (or with shallow imitations of voice). What's contested is the source of voice. The romantic position: voice is essential, lived, possibly soul-deep, impossible to fake. The engineering position (Deen's): voice is configuration across measurable axes.
Both positions agree AI is bad at voice. They disagree about why. The romantic answer: AI lacks soul. The engineering answer: AI lacks the capacity to modulate across multiple axes simultaneously in ways that respond to specific material. The engineering answer makes a falsifiable prediction: as models gain capacity to modulate, AI voice will improve. The romantic answer doesn't.
Deen's pragmatic compromise: "I think it's possible to be both mystical and analytical about voice."1 Some sentences come accidentally. But you can use analysis to figure out the dull patches of your draft and keep riffing until something feels right. The framework is the analytical layer that supplements (not replaces) intuition.
Deen's final framing: "voice is the projection of an organic personality. This means that it's sensitive, highly attuned to the material at hand. It's not just the posturing of a literary style, but it's a profound thoughtfulness around language, an ability to wield all of literature's tools to put into words how you actually feel about the object in focus."1
Two operating words: organic and attuned. Voice that's locked into one configuration regardless of material is not organic — it's mechanical. Voice that bends to the material at hand is organic. The bending is the mark of personality.
The trust function: "Organic voice earns the trust of the reader, because they can immediately smell if you're rehashing canned phrases to seem funny and deep, or if you're actually inviting them into an approximation of your consciousness."1 The reader's trust signal is whether the writer is responsive to material or just performing a brand.
Brand-voice: Same configuration across every essay. The reader trained to expect signature; signature ossifies; the writer can't evolve. The fix: deliberately deploy your axis-opposites on at least one essay per year.
Adjective-driven voice: The writer thinks their voice is "casual" or "intense" and over-pulls on whichever axis they think the adjective maps to. The fix: drop the adjective. Engineer the nine axes individually.
No friction-compensation: The prose is correct but exhausting to read because the writer never deploys Imagery, Rhythm, or Sound to compensate for reading's cognitive labor. The fix: read for friction. Where does your reader's attention sag? Add Sound or Sight at those points.
Imitation without integration: Writer studies a master's voice and ports the surface features without understanding why they work. The result reads as parody. The fix: don't imitate features. Imitate the axis-settings and see which ones serve your own material.
Over-engineering: Every sentence has visible Sound, visible Imagery, visible Subtext, visible Tone modulation. Exhausting. The fix: vary intensity. Most sentences should be plain. The engineered ones get their punch from contrast.
Nine-axis self-mapping: Read a recent essay. For each of the nine axes (Tone register/confidence/valence/distance, Perspective, Subtext, Repetition, Rhythm, Rhyme, Imagery, Words, Motif), write down your default setting. This is your baseline voice configuration.
Find your axis-opposites: For each default, write the opposite. If you default to high register, write informal. If you default to active voice, write passive. The opposite list is your growth territory.
Pick three axes to expand on the next draft: Don't try to expand all nine at once. Pick three you've been most fixed on. Deliberately deploy their opposites at least once in the draft.
Run the read-aloud test: Voice is auditory at its base. Read the draft aloud. Notice where it sounds mechanical. Notice where it sounds alive. The difference is the configuration shift between the two.
Identify your dull patches: Deen's exact phrasing. "Use analysis to figure out the dull patches of your draft, and keep riffing in those areas until something feels right."1 The framework is most useful at the dull patches. Healthy patches don't need analysis.
Where Deen breaks sharpest with Dan Wang: Wang's voice-cultivation-through-stylistic-models page treats voice as deliberate inheritance — you copy master writers' sentences until their patterns become yours. Voice is stable, accumulated, lineage-based. "You're constructing voice by absorbing models." Deen's framework explicitly rejects the stability: "Your voice should be constantly evolving with you, not frozen into a brand." The split surfaces a real question: is voice the texture you've built (Wang) or the modulation you're doing right now (Deen)? Both might be true at different time-horizons. Wang gives you the materials voice can be made of (master sentences absorbed); Deen gives you the controls voice operates through (nine axes). Neither is sufficient alone. Wang's writer ends up with rich materials but locked configuration. Deen's writer ends up with flexible controls but no inherited material to modulate from. Together they make a more complete picture neither author owns.
Where Deen partially aligns with the show-don't-tell tradition: standard show-don't-tell pedagogy treats voice as a byproduct of concrete prose — write vividly and your voice will emerge. Deen agrees concrete prose serves voice (his Sight chapter formalizes this), but rejects the byproduct framing. Voice isn't accidental. It's the deliberate output of nine deliberately tuned axes. The split: standard pedagogy treats voice as something you don't manage; Deen treats voice as something you manage exactly. The reader who absorbs both reads through writing in two layers — the felt experience (standard pedagogy) and the engineering producing the experience (Deen).
Where Deen takes a position the vault hasn't named: most writing traditions either treat voice as ineffable (romantic) or treat it as imitable through model-study (Wang-style). Deen offers a third option: voice as engineerable through axis-tuning. This third position is novel enough to be the contribution of the Essay Architecture framework. The romantic position can't be falsified. The Wang position can be tested (copy masters → see what emerges). The Deen position can be operationalized (tune axes → measure reader response). Whether it can be empirically validated is open — but it's the most testable of the three.
Voice as a concept reaches into adjacent vault territory at three specific points.
Behavioral Mechanics — Multi-Axis Presence Engineering (Hughes BOM): Hughes's behavior-ops training teaches operators to deliberately modulate across multiple presence axes simultaneously — body posture, tonal register, breathing pace, eye-contact density — to produce specific effects on a target. The structural identity to Deen's voice-as-9-axis-engineering is exact. The split: Hughes operationalizes presence for influence; Deen operationalizes voice for prose communication. Same underlying mechanism: human-readable identity is the output of simultaneous modulation across multiple dimensions, and intentional control of those dimensions produces intentional identity. What this reveals: the engineering of voice in writing and the engineering of presence in person-to-person operations are versions of the same skill. A writer who masters Deen's framework is implicitly trained in the structural literacy that behavioral-mechanics formalizes for embodied operations.
Eastern Spirituality — Anatta / Non-Self Doctrine: Buddhist anatta denies that there is an essential, stable self underneath the changing aggregates (form, sensation, perception, mental formations, consciousness). The self is the current configuration of the aggregates; change the aggregates, change the self. Deen's voice-is-not-a-thumbprint position runs the same metaphysics in the writing-craft register. There is no essential voice underneath the configurations. Voice is the configuration. The split: anatta is metaphysical; Deen's claim is operational. The structural parallel: both reject essence in favor of configuration. What this reveals: Deen's anti-romantic-voice position has serious metaphysical company. Some of the deepest contemplative traditions converge on the same anti-essentialism that Deen applies to literary identity. Voice-as-configuration is not just craft theory — it's a metaphysical commitment with two-thousand-year-old precedent.
Psychology — Internal Family Systems / Multiplicity (IFS): IFS treats the psyche as a system of parts, each with its own perspective, tone, and concerns. Health is the ability to access different parts as the situation calls for. Deen's modulation-across-nine-axes is doing the same work for the writer's voice — accessing different settings as the material calls for. The structural parallel: well-functioning identity is the capacity to rotate through internal positions rather than being stuck in one. The writer with locked voice may be displaying, in the medium of prose, the same psychological fixity that IFS treats clinically when it shows up as locked self-states. Voice-development isn't only craft. It's a form of psychological flexibility-practice. The writer who learns to modulate is, structurally, doing the same work as the IFS-client learning to access multiple parts.
Creative Practice — Voice Cultivation Through Stylistic Models (Wang): Wang and Deen produce the most productive vault tension on voice. Wang locates voice in stable inherited material (you study masters; their patterns become yours). Deen locates voice in live-time modulation (configuration shifts as material demands). Both reject the romantic frame. They split on whether voice has stable identity or is constantly evolving. The handshake suggests a synthesis neither author articulates: voice operates at multiple time-scales. Wang's frame describes long-time identity (the stable texture across years). Deen's frame describes short-time variation (per-sentence configuration). The practitioner ends up with both. Wang gives the materials; Deen gives the controls. Neither alone is sufficient. The integration is the craft.
Synthesis across the four connections: each domain has identified the same fact — operational decomposition of identity-into-axes is required to make identity engineerable. Hughes operationalizes presence into axes for influence. Anatta tradition deconstructs self into aggregates for liberation. IFS decomposes psyche into parts for therapy. Wang names voice's components (model-inheritance + life-shape). Deen names voice's nine axes for engineering. The deeper claim: identity is configuration, not essence, regardless of which domain you approach the question from. Voice that feels mystical is voice that hasn't yet been operationally decomposed. Voice-as-engineerable is the contemporary frontier; voice-as-thumbprint is the romantic position the AI period is rendering untenable.
The Sharpest Implication: If voice is configuration rather than essence, then the writer's identity as a writer is radically less fixed than the romantic frame assumes. You are not stuck with the voice you currently have. You can — deliberately — become a different writer by changing axis settings. This is uncomfortable because most writers have built their public identity (and often their sense of artistic self) on a particular voice. The implication: the writer who can evolve voice deliberately has a competitive and creative advantage over the writer who treats voice as fixed identity to be preserved. In a writing economy increasingly threatened by AI, the moat is not your particular voice — it's your capacity to modulate. The writer who can change configurations as the material changes is harder to replace than the writer who delivers consistent brand-voice essay after essay. Brand-voice is the most replicable kind of voice for AI to imitate.
Generative Questions: