Two characters share the stage. One of them speaks. Watch what happens to the audibility of the speech in the audience's reception of the scene.
In the simplest case, the speaker addresses the other character and both characters hear the words. The audience hears them too. This is ordinary dialogue, not one of the framework's named modes.
But the speaker can also operate in modes where the audibility-rules are different. NS XXVI.82-91 names four such modes — speech-configurations defined by who hears what.1
In Speaking to the Sky (ākāśa-bhāṣita), the speaker addresses someone not present on stage — an absent person, a distant figure, an imaginary interlocutor. The other characters on stage do not engage; the audience hears the speaker's side of the conversation; the audience constructs the absent interlocutor from the speaker's responses. NS XXVI.83-85: "Addressing someone staying at a distance or not appearing in person or indirectly addressing to someone who is not close by... presents the substance of a dialogue by means of replies related to various imaginary questions which may arise out of the play."2
In Speaking Aside (ātmagata), the speaker speaks to themselves audibly to the audience. The other characters on stage are supposed not to hear. NS XXVI.85-86: "When overwhelmed with excessive joy, intoxication, madness, fit of passion, repugnance, fear, astonishment, anger and sorrow one speaks out words which are in one's mind, it is called Speaking Aside."3 The speaker is in a body-state intense enough that words leak out involuntarily; the convention is that other characters do not hear what should be unspeakable.
In Concealed Speaking (apavāritaka), the speaker conceals speech from one specific other character while another on-stage character hears. NS XXVI.86 names this briefly — "Concealed Speaking is related to secrecy."4 The audience and one specific character hear; another specific character does not.
In Private Personal Address (janāntika), one character addresses one other character privately while a third character (who is physically close by) is supposed not to hear. NS XXVI.87-88: "When out of necessity persons standing close by are supposed not to hear what is spoken to someone else, this constitutes Private Personal Address."5 The convention is that proximity does not automatically produce hearing; the staging-cue (one speaker leaning to one specific other, voice lowered) tells the audience that the third character does not hear.
The framework treats audibility-configuration as operationally precise rather than incidental. Each mode has a defined who-hears-what structure, defined staging-cues (the body turns away, the body leans toward, the voice lowers, the eyes fix on something invisible), and defined dramatic-functional uses.
Modern dramatic-craft preserves all four conventions operationally. The aside (Speaking Aside / ātmagata) is the Shakespearean soliloquy and the Renaissance-stage convention. The soliloquy-to-the-sky (Speaking to the Sky / ākāśa-bhāṣita) is the prayer-to-an-absent-figure and the address-to-an-imagined-other. The whispered-aside (Private Personal Address / janāntika) is the staged-confidence in modern theatre and certain cinematic conventions. Concealed Speaking (apavāritaka) is the staged-double-meaning where some characters hear what others miss.
The framework's claim is that these four are operationally distinct mechanisms with different staging-cues, different audience-reception-rules, and different dramatic-functional purposes. Modern drama deploys them but rarely articulates them as a four-fold catalog. Naming the four explicitly makes them deployable rather than dependent on tradition-transmitted convention.
The four-modes framework speaks directly to several modern crafts:
As stage-acting convention library. Modern theatre preserves these conventions operationally (the aside, the soliloquy, the whispered confidence, the prayer-to-the-absent). The framework names them explicitly as a four-fold catalog, which gives writers and directors deployable craft-vocabulary.
As cinema-translation reference. Cinema replaces these conventions with technical equivalents. Voiceover replaces ātmagata (the audience hears the character's interior speech while the on-screen body says nothing). Cutaway-to-listener replaces janāntika (the camera-cut shows what the close-by character does not hear). Cinematic isolation-shot replaces ākāśa-bhāṣita (the close-up that excludes the on-screen world while the character addresses an absent figure). The framework names what cinema is translating into camera-conventions.
As podcast/audio-monologue craft. The ākāśa-bhāṣita mode is the structural foundation of podcast monologue. The host addresses an absent "you"; the listener constructs the implied audience-party from the host's responses. Modern podcast monologue uses the convention instinctively; explicit naming clarifies the craft and gives the host deliberate-deployment vocabulary.
As subtext-writing tool. The four modes are subtext-architectures with operational rules. A scene that requires subtext can deploy janāntika (private personal address) for shared-confidence between two characters; apavāritaka (concealed speaking) for selective-hearing across the cast; ātmagata (speaking aside) for character-interior leakage. The framework's vocabulary gives subtext-writing more precision than the generic "subtext" category modern craft typically deploys.
Four-as-canonical: NS gives four. Other dramatic traditions name fewer or more speech-modes. The four are working-minimum for classical Sanskrit dramatic representation; modern cinema and recorded media have added technical-conventions (the voiceover, the cutaway, the close-up isolation) that may constitute additional modes or may map onto the four.
Convention-dependence: the modes work because the audience knows the convention. Audiences untrained in the conventions cannot read them — they see two characters on stage and assume both hear what is spoken, regardless of the staging-cues the framework prescribes. Modern adaptation depends on whether the conventions are still operational in audience-expectation; some are (the soliloquy, the voiceover), others have weakened (the staged-whispered-confidence does not always read as inaudible-to-the-close-by-third-character in modern naturalistic conventions).
Live-vs-recorded translation: the modes were developed for live theatre. Recorded media translates them through different technical mechanisms — the modes operate in cinema and audio but through camera-cuts, voice-over, audio-mixing rather than through stage-conventions. The framework's principle (four distinct audibility-configurations are operationally meaningful) survives translation; the specific staging-cues do not.
Bharata is the single primary source for the four modes, and the productive tension is between speaker-audience configuration specificity (each mode names a precise speaker-listener pattern) and staging-convention dependence (audiences must know the conventions). The framework provides operational specificity that requires audience-training to deploy.
This is consistent with NS's broader pattern. The framework operates on the assumption that the audience knows the dramatic conventions — that the audience is competent to read the staging-cues and apply the corresponding audibility-rules. This assumption is partly cultural-historical (Sanskrit dramatic-tradition audiences were trained on the conventions across centuries) and partly cognitive (the underlying convention-recognition capacity is human-universal even when the specific conventions are culturally trained).
Modern adaptation can take the four as available conventions that can be re-established even if they have weakened in some audience-contexts. Cinema has done this for ātmagata (voiceover is now universally understood); theatre has preserved the soliloquy more strongly than the private-personal-address. The conventions can be cultivated in audiences through consistent deployment; the framework's claim is that the conventions are operationally available if practitioners commit to deploying them clearly.
The four-modes framework is a creative-practice page, but the underlying claim — that speech-mode is operationally defined by audience-listener configuration — runs into adjacent vault domains.
behavioral-mechanics: Selective Honesty and Micro-Revelations — influence-mechanics describes how operators manage what is revealed to whom through specific speech-configurations. The structural parallel to the four dramatic-speech modes is exact: both frameworks treat audience-listener differentiation as operationally precise rather than incidental. NS provides the dramatic-context configuration-vocabulary; influence-mechanics provides the operational-context configuration-vocabulary. Combined: skilled communication (whether dramatic or operational) requires explicit attention to who can hear what — and the framework's four modes provide working categories. The handshake exposes that audibility-management is a craft-resource that operates across dramatic-craft and influence-operations. The dramatist managing audience-and-character-audibility through janāntika and ātmagata is doing the same operation the operator managing selective-revelation-among-multiple-audiences is doing — different deployment, same underlying audibility-architecture craft.
psychology: The Relationship Between Storage and Retrieval in Memory — memory-and-discourse research shows that internal speech and external speech operate on different cognitive substrates. The structural parallel: the four dramatic-speech modes operationally model the internal-external distinction in performance. Ātmagata (speaking-to-oneself audibly) is the staged version of internal speech; ākāśa-bhāṣita (speaking-to-imagined-other) is the staged version of imagined-dialogue; the others operate at varying internal-external positions. The framework names what cognitive psychology has only recently studied. The handshake reveals that the dramatic tradition observed the internal-external distinction in cognitive speech and codified it as four distinct speech-modes; modern cognitive-research has independently confirmed that internal-and-external speech operate on different cognitive substrates. The framework's two-millennium-old categorization reaches what modern cognitive-science has had to discover through experimental research.
The four modes are operationally preserved in live theatre but compressed and translated in cinema and recorded media. Modern adaptation could recover the explicit four-mode vocabulary to upgrade craft across forms — particularly in podcast monologue and direct-address video, where the ākāśa-bhāṣita mode operates centrally without typically being articulated as a specific named technique.
Generative Questions