Source files: earnings-call-voice-analysis-2026-raw.md
Source status: peer-reviewed papers and primary research pages prioritized; working papers, vendor claims, and inferences are labeled. This note does not identify a profitable live hedge-fund voice factor.
Executive answer
The defensible thesis is not “AI can tell when an executive is lying.” It is that a speaker-normalized, topic-conditioned representation of how management speaks may add information to what management says—particularly in unscripted Q&A and when voice and transcript disagree.
The evidence is strongest for four narrower claims:
- Vocal affect and delivery contain information incremental to reported earnings and transcript words.
- Some acoustic embeddings have shown predictive relationships with future returns or volatility in research samples.
- Voice–text interaction may help separate routine positive or negative language from language delivered with unusual strain or confidence.
- Vocal or multimodal features can support fraud-risk triage, but they do not establish a reliable automated lie detector.
What researchers are measuring
| Layer | Example features | Interpretation to test |
|---|---|---|
| Acoustic | Pitch, energy, jitter, shimmer, harmonicity, speech rate, pauses, breathing, spectral features | Change in delivery relative to the speaker’s own baseline |
| Affective | Fear, anger, excitement, arousal, valence, dominance, nervousness | Soft information about perceived conviction or concern; not ground-truth emotion |
| Linguistic delivery | Fillers, repetitions, repairs, hedging, response latency, complexity | Spontaneity, cognitive load, evasiveness, or communication difficulty |
| Dialogue | Analyst question type, interruption, follow-up pressure, response directness, turn timing | Whether the executive changes under scrutiny |
| Cross-modal | Positive words plus strain; negative words plus calmness; topic-specific disagreement | Incremental information beyond text-only sentiment |
| Delivery quality | Comprehensibility, steadiness, clarity, pace, balance | Whether the message is easy for a listener to process |
Core research findings
Vocal affect and future fundamentals
Mayew and Venkatachalam’s Journal of Finance study analyzes earnings-call audio with vocal-emotion software. It finds that positive and negative affect are associated with contemporaneous returns and future unexpected earnings, after controlling for quantitative earnings information and linguistic content. The relation to future earnings extends over the two subsequent quarters, with the strongest effects when managers face analyst scrutiny.
That is a forecasting horizon, not proof of a two-quarter trading lead. The market reaction studied around the call is much shorter.
Fear and anger are context-dependent
Sul, Wang, and Zhu’s 2025 working paper distinguishes vocal fear from vocal anger. In the paper’s S&P 500 sample, fear is associated with negative investor reactions and later underperformance, particularly when it conflicts with positive earnings news. Anger is associated with positive reactions and future growth in contexts where earnings news is strong.
The implication is methodological: a single positive/negative voice score is likely to discard context. The same acoustic activation can mean concern, conviction, frustration, or emphasis depending on the news, question, executive, and topic.
Delivery quality is not the same as emotion
The 2025 Vocal delivery quality in earnings conference calls paper uses a deep-learning measure of acoustic comprehensibility. It reports real-time market reactions and shows that delivery quality varies with the kind of news being communicated.
This suggests a separate factor family: not “does the executive sound happy?” but “how difficult is the information to receive and process?” It can be tested as a risk or attention variable rather than interpreted psychologically.
Deep embeddings are promising but model-sensitive
The 2025 Netspar industry paper compares TRILLsson and W2V2 embeddings on short CEO segments from opening remarks and Q&A. The published summary reports W2V2 predictions above chance at 54–59%, while TRILLsson does not show the same predictive result. Results vary by model, method, and call segment.
This is evidence that representation choice matters. It is not evidence that “voice models” as a category generate a stable, cost-adjusted portfolio.
Voice–text interaction may be the practical route
Pope’s 2026 working paper studies 41,395 Russell 3000 call observations from 2020–2025. Its abstract reports that, within the lowest text-sentiment group, high vocal strain was associated with approximately 103 basis points of lower excess return over 20 trading days relative to low strain. In the highest text-sentiment group, high vocal valence added approximately 39 basis points over 10 trading days.
This is directly relevant to a trading design, but the source is employer-funded and the author discloses a financial interest in the supplying firm. The numbers should be treated as a replication target, not an established market fact.
Deception and fraud: useful as triage, not a lie detector
Burgoon et al. (2016) found that restatement-related utterances differed on vocal and linguistic dimensions, especially when comparing prepared remarks and unscripted responses. A 2024 multimodal fraud-detection paper combines verbal, vocal, and analyst–manager dialogue features and reports an 8.64% improvement over its best baseline on a 4,203-call research sample.
Neither study establishes a forward-return lead time. Fraud labels often become available only after restatements, investigations, or enforcement actions. They also measure “associated with later misreporting,” not necessarily “the executive was lying in this sentence.”
A public fund implementation: Plato’s Q&A-evasion triage
Plato Investment Management has described a related, text-first workflow. In a July 2026 article, Plato quantitative-research analyst Marcus Howes says the team refined an earnings-call “evasion” detector over approximately three years. The system isolates unscripted Q&A and evaluates three features: whether an answer stays on topic, the answer-length/question-length ratio, and the share of future-tense language. Plato says an LLM is used specifically for topicality; a composite flag is then created when the features align.
This is a useful implementation comparator because it makes the decision boundary visible. The output is a review trigger inside a broader red-flag process, not a verdict that a speaker lied. The same article gives examples in which long answers or reframing could have benign explanations, and does not claim that the classifier establishes deception. Plato’s team page identifies Howes as an Associate Analyst working with machine learning, neural networks, NLP, and LLMs; Plato’s investment-process page separately describes NLP-derived tone from roughly 25,000 earnings calls per year. These are firm and practitioner disclosures, not an independent audit of the model, its labels, its false-positive rate, or its trading results.
The same team page exposes a separate automation owner. Senior Quantitative Analyst Wilson Thong is described as overseeing automation across production, compliance, and reporting, and as the creator of PRISM, an internal business-intelligence platform covering factors, the proprietary red-flags model, portfolio exposures, attribution, stress tests, and scenarios. It also lists Senior Portfolio Manager Chanel Stuart-Findlay as having previously published work on incorporating machine learning and NLP into investment processes. These are useful personnel and workflow signals, but the page does not connect either person to a particular model, permission set, or automatic trade decision.
The features are semantic and dialogue-based rather than acoustic. They should therefore be kept separate from pitch, energy, pause, and other voice measures. The public record also does not establish that the signal is used for automatic position changes, nor does it provide a reliable forward lead time. The most defensible replication target is a point-in-time Q&A review queue whose output is “topic-specific response anomaly,” followed by human fundamental diligence.
The broader deception literature is a material counterweight. The Annual Review of Psychology review concludes that known nonverbal cues to deception are generally faint and unreliable. A production system should therefore output “topic-specific delivery anomaly requiring diligence,” not “lie detected.”
Lead-time map
| Use case | Evidence-supported timing | What it means |
|---|---|---|
| Immediate market reaction | Same call to roughly two trading days | Mayew and Venkatachalam study short-window abnormal returns; delivery-quality research reports real-time reaction. |
| Post-earnings drift | 10–30 trading days | Reported in the 2026 practitioner voice–text study; requires independent replication and cost analysis. |
| Fundamental information | One to two subsequent quarters | Future unexpected earnings relation in the Journal of Finance study. This is not a guaranteed monetization window. |
| Volatility or risk | Call/event horizon | Multimodal models predict risk or volatility, but the exact useful horizon is model- and dataset-specific. |
| Fraud/deception | No established lead time | Current research supports classification or triage experiments, not a reliable advance warning interval. |
Potential monetization designs
1. Conditional post-earnings selection
Use earnings surprise and transcript sentiment as the base event. Add a voice residual:
voice residual = current delivery vector
− expected delivery for this executive, role, topic, and channel
Trade or size only when the residual adds information after controlling for sector, size, liquidity, volatility, prior returns, analyst coverage, and guidance surprise.
2. Analyst workflow and estimate revisions
Calls with large topic-specific deviations can be routed to fundamental analysts. The output is a ranked review queue, not an automatic buy or short instruction. A useful label is: “positive language, elevated strain around gross margin,” not “management deception.”
3. Risk and short-book overlay
Voice–text disagreement can be used to reduce position size, increase required diligence, or review short exposure. It may have value even if it does not predict direction, because it can signal uncertainty or a wider distribution of outcomes.
4. Volatility and options research
Delivery instability, question pressure, and low comprehensibility can be tested against realized volatility, implied volatility, volume, and spreads. This should be evaluated as a volatility forecast, not automatically converted into a directional trade.
5. Data licensing
Commercial vendors publicly market speaker-normalized audio and behavioral features. Speech Craft Analytics describes pitch, jitter, shimmer, energy, pauses, speech rate, disfluencies, arousal, valence, and event-study horizons. FactSquared describes transcript-plus-audio stress and speaking-pattern analysis for investment users.
A separate Quartr customer story about Kepler exposes a vendor-reported architecture adjacent to this data problem: Claude interprets the analyst’s question, deterministic code retrieves the underlying financial data and runs calculations, and a financial ontology connects the language layer to companies, filings, line items, and events. Kepler’s output is described as traceable to phrase-level primary-source material. The page says the platform serves hedge-fund, private-equity, and bank analysts and uses Quartr earnings-call and IR materials across 15,000+ companies and 65 markets. These are customer-story and vendor claims, not confirmation of a named fund’s adoption, a partnership with a fund, or predictive performance.
The public record does not name a complete hedge-fund client list or provide independently audited portfolio results for these vendors.
Datasets and benchmarks
There is no single public, universally labeled archive of every earnings call with high-quality emotion, deception, speaker, and return labels. The available resources cover different tasks:
| Resource | What it provides | What it does not provide |
|---|---|---|
| Qin & Yang dataset | Audio and transcript data associated with multimodal volatility research | Universal coverage or standardized emotion labels |
| Earnings25 | 500 hours of finance speech with aligned transcripts and speaker/industry metadata for ASR | Voice sentiment, deception, or return labels |
| FinCall-Surprise | Multimodal earnings-call benchmark for earnings-surprise prediction | A validated acoustic factor |
| MiMIC | Transcripts plus images and tables from Indian earnings presentations | A broad voice-emotion benchmark |
| 4,203-call fraud sample | Multimodal fraud classification research sample | Evidence that the complete labeled corpus is publicly downloadable |
The practical research asset is usually a licensed or assembled archive of audio and transcripts, joined to point-in-time prices, estimates, guidance, speaker identity, call metadata, and later fundamental outcomes.
Required model and backtest controls
- Diarize speakers and isolate CEO/CFO responses from operator and analyst audio.
- Use executive-specific baselines; raw pitch or speech rate is not comparable across people.
- Separate prepared remarks from Q&A.
- Preserve filler words, repetitions, repairs, pauses, and timestamps; commercial transcripts may remove them.
- Normalize for microphone, codec, conference provider, language, accent, illness, and call time.
- Freeze features at the moment the call becomes tradable; avoid later transcript corrections or revised labels.
- Use temporal walk-forward validation and purged event splits, not random row splits.
- Test incremental value against transcript-only, earnings-surprise, price-reaction, and volatility baselines.
- Include spreads, borrow, market impact, turnover, and audio-ingestion latency.
- Report calibration, coverage, false-positive rate, and performance by sector and speaker—not only aggregate accuracy.
Evidence boundaries
The public research supports a testable proposition: management voice and delivery may contain incremental soft information. It does not establish that any particular emotion is a universal trading signal, that voice detects deception reliably, that vendor-reported return differences survive costs, or that any named hedge fund operates a live voice factor.
The next decisive experiment is an independent point-in-time replication of the voice–text interaction result, using speaker-normalized Q&A features and a pre-registered 1-, 5-, 10-, 20-, and 30-trading-day event study.