A Folklore: Are Female Voices Really More Effective?

Conversational AI vendors often claim that female-presenting voices engage consumers better and resolve more accounts. Neither the published literature nor live deployments support the claim. This preliminary research note sets out the evidence, and why Krew will not choose a voice to trade on a gender stereotype even if it did work.

Preliminary Research Note · September 2026

These findings do not justify a universal female-voice default for collections.

Voice gender should be treated as a configurable parameter and evaluated within each portfolio against payment outcomes, rather than selected on the assumption that one gender is inherently more effective.

The claim

A recurring claim in the conversational AI market is that female-presenting synthetic voices engage consumers more effectively than male-presenting voices and, in its strongest form, resolve more accounts.

The published literature reviewed here does not establish a general female-voice advantage. Nor did our review identify a peer-reviewed study directly testing synthetic voice gender against collections outcomes.

What we found

The reviewed evidence supports three narrower conclusions: gender-stereotyped trait attribution does not establish a consistent advantage in trust or persuasion; responses to voice gender can depend on perceived role and context; and perceived age, tone, and listener characteristics complicate any rule based on gender alone.

In the deployments reviewed here, where each client assigned voice at random within every strategy, we find no consistent female-voice advantage: the difference is within a second and a fraction of a point of zero, and where it is large it reverses sign between portfolios running the same voices in the same months.

A literature review

Where voice gender moves outcomes, it does so as an interaction, not a main effect

  1. 1. Trust effects depend on context and role expectations

    Zhang, Yang, and Robert (2025) ran a randomized study and found that similarity between listener gender and voice gender raised affective trust, but that this effect held only when the voice was congruent with stereotypical role expectations.

    Jeon (2024) provides a direct reversal of the industry claim in a commercial setting. Across brand-concept conditions, male-presenting AI agents produced significantly higher trust than female-presenting agents for functional brands (M = 4.16 vs. 3.66), while for experiential brands the difference was not statistically significant.

    Jeon’s finding suggests that male voices might elicit greater trust than female voices in collections, insofar as the advantage observed for male agents in functional-brand contexts extends to repayment conversations.

  2. 2. Gender alone is an incomplete basis for voice selection

    Pias et al. (2024) examined persuasiveness in voice-based product recommendation and found that positive and neutral tones outperformed negative ones, and that the most persuasive voices were those perceived as middle-aged male or younger female. Gender did not order the conditions; it interacted with perceived age and tone in a way that no single-variable rule captures. Tolmeijer et al. (2021) similarly treated pitch as an independent manipulation rather than a proxy for gender.

    For deployment purposes this matters because tone, pace, pitch and turn-taking behavior are all tunable within any voice, whereas “select the female voice” is a single binary choice that discards the dimensions carrying the variance.

  3. 3. Listener heterogeneity limits generalization

    Moradbakhti, Schreibelmayr, and Mara (2022) found that male participants reported significantly lower autonomy satisfaction when interacting with high-agency female-presenting assistants than with low-agency ones, and that this reduced satisfaction predicted lower intention to use the assistant; female participants showed no comparable differentiation across conditions. An effect measured on one population therefore need not survive transfer to another.

    A collections portfolio is not a homogeneous population: stage of delinquency, balance, product, and demographics vary within a single book, and the literature gives no reason to expect a uniform voice effect across those segments.

LiteratureJeon (2024)

Trust in an AI agent, by agent gender and brand type

Mean trust rating; higher is more trust. n.s. = not significant.

Female agentMale agent
Voice gender: no consistent effect

No consistent voice-gender effect on engagement or conversion

Adjusted for direction and portfolio, the two voices differ in engaged time by 0.1 seconds (95% CI −0.4 to +0.6). On promise to pay, the outcome a collections operation is actually run on, they differ by 0.1 percentage points (−0.2 to −0.0), a tenth of a point favoring the male voice, with an interval that reaches zero.

That pooled figure is the headline, not the finding. Forty-four thousand live conversations, voice randomly assigned within every strategy, inbound and outbound, two portfolios, one voice per gender, channel and book mix removed, put the average difference within a second and a fraction of a point of zero.

The pooled figure averages strata that disagree. The four direction × portfolio strata are heterogeneous on every outcome, and the inverse-variance weights are dominated by the two outbound strata, which hold 39,000 of the 44,000 conversations. The inbound Book C stratum, where the female voice holds the line roughly 37 seconds longer, carries almost no weight. The stratum-level results, not the pooled figure, are the substantive finding of this note.

A natural defense of the claim is that averages conceal what happens once a conversation gets going. On the pooled numbers, they do not. Within long conversations the female voice holds 87.9 seconds against 90.7 for the male voice, and promise to pay is nearly identical.

The female voice shows no consistent advantage on any outcome; where it leads in one stratum, it trails in another.

Primary result44,119 conversations

Adjusted difference, female voice minus male voice

Stratified by direction × portfolio, 95% CI. Zero means no difference.

Engaged time

+0.1s

95% CI −0.4 to +0.6

Promise to pay

−0.1pts

95% CI −0.2 to −0.0

Do listener demographics matter?

The similarity mechanism is, at most, weak

The most credible mechanism by which a voice-gender advantage could exist is similarity between speaker and listener. We find at most weak evidence of it.

Among female consumers the female-minus-male difference in engaged time is −0.4 seconds (−1.2 to +0.3). Among male consumers it is −1.4 seconds (−2.2 to −0.6). The difference between those two differences is +1.0 second (−0.2 to +2.1), which does not exclude zero. These contrasts use the 72% of calls with a classifiable first name.

The pattern runs in the direction the similarity hypothesis predicts: male consumers engage 1.4 seconds less with the female voice, a small but non-zero gap, while among female consumers the difference is near zero. But the test the similarity mechanism requires is the contrast between those two gaps, and that interaction is small and not statistically distinguishable from zero (p ≈ .08). On long-conversation share the homogeneity test across consumer gender returns p = .78.

This is absence of evidence rather than evidence of absence: if a similarity effect exists here, it is on the order of a second, and it does not rescue a general female-voice advantage.

Listener demographics72% of calls classifiable

Engaged time: female voice minus male voice, by consumer gender

Seconds, 95% CI. Consumer gender inferred from first name, in aggregate only.

Between portfolios

Where a difference is large, it does not transfer between portfolios

This is the finding with the most practical consequence for anyone evaluating a vendor benchmark, and it is not a null.

Split by portfolio, on outbound calls: on producing a long conversation, the female-minus-male difference is −8.0 points in Book C against +0.6 in Book D (Cochran’s Q = 259.8, p < 10−15). On whether the consumer engages at all, +0.8 against +4.5 (Q = 33.4, p = 8 × 10−9). Across the four direction × portfolio strata, the voice difference is heterogeneous on every outcome measured.

An eight-point deficit in one portfolio and a slight lead in the other, with the same vendor, same voices, same months, and voice randomized within every strategy, is a reversal, not noise, and because assignment is random within strategy it is not the voices running on different scripts or campaigns. Whatever is being measured is not a property of the voice that travels with the voice. It is a property of a configuration in a context, and moving the voice to another book does not move the result with it.

This is the point at which the industry claim runs out, and it runs out on design rather than on data. “Female voices perform better in collections” is a general claim, and a general claim requires a design capable of testing generality: more than one portfolio, so that a book-specific effect can be seen to reverse; the script held constant across voices, so that the voice is not simply a label on a different conversation; channel stratified and disclosed; and more than one voice per gender. A benchmark that satisfies none of these has not tested the claim; it has run a case study on one configuration and reported it as a law. The comparison in this note is randomized within every strategy by each client’s own configuration, which holds script and calendar constant by construction, but it tests one recording per gender, and we accordingly report it at the stratum level rather than resting on the pooled figure.

The claim as it circulates has not, to our knowledge, been accompanied by any of these. And our own data shows why that matters: every one of those shortcuts produces the folklore’s result for free. Pool channels without stratifying and the female voice leads by 1.9 seconds. Measure on a book that happens to resemble our Book D and the female voice is ahead on whether the consumer engages at all and on producing a long conversation. Let the voices run different scripts and the voice absorbs the script’s effect. Test one recording per gender and the result is a fact about two files. None of these requires an error in measurement. They require only that the design be too narrow for the conclusion drawn from it, which is the ordinary condition of a vendor benchmark.

Grant every benefit of the doubt. Assume the comparison was randomized, the intervals honest, the sample large. A single-portfolio result still cannot support a claim about voices in general, because in this data the identical randomized comparison changes sign between two books running concurrently. The most generous reading of the folklore is that it is true somewhere, as a real effect of a particular voice, on a particular script, on a particular line, against a particular book, and has been sold as true everywhere.

A benchmark from another portfolio is not evidence about yours. Not at any sample size, and not at any p-value.

Between portfoliosoutbound, same voices, same months

The same comparison reverses between portfolios

Female minus male voice, percentage points, voice assigned at random within strategy.

Book CBook D
Sources of variation

Who’s being called, and with what AI, better explains heterogeneity

Heterogeneity does not stop at the portfolio boundary. Holding the book fixed, the difference fails a homogeneity test at p < .01 on a substantial share of outcome-by-segment checks across delinquency stage, balance, and region. Within Book C, whether the consumer engages at all runs from −6.5 points to +6.8 points across delinquency stages (p < 10−15), a thirteen-point swing inside a single book.

Delinquency stage, which the industry claim never mentions, moves the measured difference by thirteen points. Consumer gender, the one characteristic the claim is about, moves it by about a second. Practitioners asking which voice suits which segment are asking a sound question about the wrong variable.

Holding the voice constant, with the same female voice, on the same inbound line, with the same model, and the same safeguards, engaged time runs from 54.4 seconds (48.3–60.5) under one persona to 24.3 seconds (23.9–24.8) under another, with three further personas distributed between them. The personas serve different client portfolios, so this bounds the combined effect of script and caller population rather than script alone. Thirty seconds of difference from the conversation alone, against a pooled voice difference of a tenth of a second.

The variable that reliably moves the outcome, and that a deployment genuinely controls, is what the agent says and how the flow is built, not who appears to be saying it. A team debating voice gender is optimizing the smallest lever available to it while the largest sits untouched.

Sources of variationsame data, same window

What actually moves the outcome

Range of the measured difference attributable to each factor.

Engaged time, seconds

Persona and caller population, same female voice0.0 s
Consumer gender (listener–voice similarity)0.0 s
Voice gender, adjusted0.0 s

Whether the consumer engages at all, percentage points

Delinquency stage, within Book C0.0 pts
Voice gender, adjusted0.0 pts
What a credible comparison requires

Treat voice gender as a parameter, not a default

Voice gender should be treated as a configurable parameter with a portfolio-specific setting, not as a default inherited from the vendor. Before any benchmark is treated as evidence about a voice, it should meet the conditions below. They are the questions a buyer can put to any vendor making the claim, and the standard this note holds itself to. Whatever the design, every voice in the comparison is held to the same consumer standard: the same script, offers, disclosures, and rules.

  1. 01

    Assignment at the account level

    Voice is assigned per account, not per call, so that repeat contacts hold voice constant and the unit of analysis matches the unit of outcome. Assignment is balanced on the account characteristics that drive outcomes: asset type and balance band at minimum. If the allocation ratio varies by strategy, the analysis is run within strategy or weighted for it.

  2. 02

    Script and calendar held constant

    This is the condition most often missing. If each voice runs the flow written for it, or the two voices are live in different months, the comparison measures a bundle and cannot be attributed to the voice. Random assignment within each strategy, as the clients in this note configured, satisfies it by construction. One script, two voices, same window, or the result is uninterpretable as a comparison of voices, however large it is.

  3. 03

    Channel stratified and disclosed

    An unadjusted comparison hands the female voice a 1.9-second lead purely because it carried more inbound calls. A benchmark that pools inbound and outbound without stratifying, or that does not disclose what share of each voice’s calls were inbound, cannot be distinguished from a measurement of the phone line rather than the voice, and until it is, it should not be treated as evidence about the voice.

  4. 04

    Aggregation fixed in advance

    The strata and the window are set before outcomes are examined and published alongside the result. An unstated aggregation choice is enough to produce whichever finding is wanted without a single false number.

  5. 05

    Payment outcomes, not proxies

    Right-party contact, promise-to-pay, kept-payment conversion, and time-to-first-payment are the outcomes of interest. Hang-up rate, escalation requests, and complaint rate are reported alongside them as guardrails, since a voice that improves contact while raising complaints is not an improvement.

  6. 06

    A sample large enough to see the effect

    Effects in this literature are small and conditional, and collections outcomes are low-base-rate. At a 1% baseline, detecting a 20% relative lift requires approximately 42,700 calls per arm at 80% power (α = .05, two-sided), and a 10% relative lift requires approximately 163,000. A comparison of a few hundred or a few thousand calls cannot distinguish a real effect from noise in either direction, which is the most common way category folklore gets manufactured.

  7. 07

    Consumer gender accounted for in the analysis

    A comparison run on a gender-skewed book is in principle confounded between voice gender and listener–voice similarity, so consumer gender is included as a covariate in the analysis before any conclusion is drawn about the voice itself. The listener-demographics result finds a small interaction in the direction the similarity hypothesis predicts, not statistically distinguishable from zero. The control remains good practice; the mechanism it guards against is at most weak here.

  8. 08

    Results reported by segment

    A pooled null can conceal offsetting segment-level effects, and a pooled positive can be driven by a single segment. This is not a theoretical caution: the difference reverses sign between portfolios, and varies by thirteen points across delinquency stages within a single portfolio. A result holds for the book it was measured on, and a changed book calls for a fresh measurement.

  9. 09

    Tone and pitch varied within voice

    Because these parameters carry effects independent of gender (Pias et al., 2024; Tolmeijer et al., 2021), a comparison that varies only gender is likely to be testing the least informative dimension available.

Scope and interpretation

The adjusted estimate is a weighted average over strata that disagree

The direction × portfolio strata are heterogeneous on every outcome, so the pooled figure is a summary across lines rather than a description of any one of them, and its interval should be read that way; the stratum-level results are the substantive finding and are reported alongside it.

A regression with additive direction and portfolio covariates returns +1.2 seconds and +0.3 points on the two primary outcomes. We report the stratified estimate as primary because an additive model assumes a common effect across strata, which the portfolio result shows does not hold; either way, both methods place the difference within a second and a fraction of a point of zero, on either side of it.

Scope. The analysis covers a four-month window, one voice per gender, and portfolios from different clients in US consumer collections. Each client assigned voice at random within every strategy, at ratios it chose (about 2:1 female-to-male on most strategies, up to 10:1 on some), so the comparison is randomized, not observational. Because the ratio varies by strategy, the direction × portfolio strata pool strategies with unequal allocations; a strategy-level stratification is the natural refinement. Engagement figures describe conversations in which the consumer spoke, a gate on which the voices differ by 1.5 points adjusted.

Scopedirection × portfolio

The adjusted estimate averages strata that disagree

Female minus male voice, engaged time in seconds, 95% CI, by stratum.

Design parameters

How the comparison was built

The voice comparison covers 44,119 conversations in which the consumer spoke, 29,554 on a female-presenting voice and 14,565 on a male-presenting voice, across both inbound and outbound lines and both portfolios in the window. Within every strategy, inbound and outbound alike, each client assigned voice at random at a ratio it chose, roughly 2:1 female-to-male on most strategies and up to 10:1 on some, which is why the two arms are unequal in size and why the ratio varies between strata. Voice labels in the strategy table were verified against the deployed configuration before analysis. The data is aggregated and anonymised, identifies no individual or client, and is used with the clients’ agreement.

Inbound and outbound conversations differ in kind: inbound callers are self-selected, stay on the line nearly twice as long, and convert at many times the rate. Because the two voices do not carry inbound in the same proportion, a naive pooled comparison would credit the female voice with the inbound channel’s advantages. The primary estimates are therefore computed within each direction × portfolio stratum and pooled with inverse-variance weights, which removes channel mix and portfolio mix from the comparison. Because the strata disagree, the pooled figure is reported for comparability with the unadjusted number; the stratum-level estimates are the substantive result. A regression with direction and portfolio covariates serves as a cross-check.

Definitions. Engaged time is seconds on the line for a conversation in which the consumer spoke. Promise to pay is a payment agreement record, a payment-commitment tag in call memory, or a payment method the consumer provided. Intervals are 95% Wilson intervals on rates and 95% normal intervals on means. Consumer gender is inferred from first name (72% classifiable), after the fact and only for the aggregate comparisons reported here; it played no part in how any call was handled.

How consumers were treated
  • Both voices ran the same script, the same offers, and the same disclosures.
  • Both voices ran under the same compliance rules: contact hours, frequency limits, and consent, enforced identically.
  • A person was available on request in every conversation, whichever voice answered.
  • No consumer was offered less, contacted more, or treated differently because of the voice assigned.
  • Nothing about a consumer was used to choose a voice; assignment was at random within each strategy.
Design parameters

Why the comparison has to be stratified by channel

The two voices do not carry inbound calls in equal proportion. Female minus male voice, 95% CI.

Adjusted = stratified by direction × portfolio, inverse-variance pooled.

Ethical considerations

A regulator or journalist asking why a particular voice was chosen for past-due accounts deserves a better answer than a vendor’s unpublished assertion.

The design decision

Selecting a voice for its perceived warmth, on the theory that it will change how people respond, is a design decision with a defensibility requirement attached, independent of whether it works. Two of the sources reviewed here make the point explicitly: West et al. (2019) treat default-female assistant design as the propagation of a service-role stereotype, and Moradbakhti et al. (2022) conclude that assistants conforming to conventional gender expectations may reinforce those expectations in users. Pias et al. (2024) raise the parallel question of the ethical limits of persuasion techniques in voice commerce.

The people on the other end of these calls are often at a difficult moment in their lives. A voice chosen to trade on a stereotype of who sounds trustworthy is a choice made about them without their knowledge, and it asks them to carry the cost of a belief the evidence does not support. Whether or not such a choice ever moved a number, it is not one a responsible operator should want to explain to the people it was used on.

Krew’s position

A conversation about money should be won on clarity, accuracy, and respect: the right information, delivered plainly, with a real path to resolution and a person available when one is needed. Voice is a configurable parameter that the client sets and can change, the mix is recorded, and outcomes are measured against the guardrails that matter to the consumer as much as to the operator, including hang-ups, escalation requests, and complaints. A voice that lifts contact while raising complaints is not an improvement, and we would not call it one.

That is also why this note exists. The claim that one voice resolves more accounts has circulated for years without published evidence. We would rather measure it, report a null, and publish the method than sell the folklore, and we hold ourselves to the same standard we ask of any benchmark: randomized, stratified, pre-specified, and disclosed.

Prepared by

Michael Goh

Michael Goh

Cofounder, Member of Technical Staff

Formerly led speech benchmarking at Artificial Analysis. Previously a FinTech VC and Bain consultant. MS Computer Science, University of Chicago; BA Economics, University of Oxford.

Rafael Khaykin

Rafael Khaykin

Cofounder, Member of Technical Staff

Formerly a research engineer at Figure Robotics, working on dexterity algorithms and latency research. Computer Science, University of British Columbia.

Alice Zhang

Alice Zhang

Head of Product Compliance, Member of Technical Staff

Formerly advised the United Nations on AI Ethics and Safety for defense industries. Previously In-House Counsel at a FinTech VC. BA (LLB) Jurisprudence, University of Oxford.

References
  1. Jeon, J.-E. (2024). The effect of AI agent gender on trust and grounding. Journal of Theoretical and Applied Electronic Commerce Research, 19(1), 692–704. https://doi.org/10.3390/jtaer19010037
  2. Moradbakhti, L., Schreibelmayr, S., & Mara, M. (2022). Do men have no need for “feminist” artificial intelligence? Agentic and gendered voice assistants in the light of basic psychological needs. Frontiers in Psychology, 13, Article 855091. https://doi.org/10.3389/fpsyg.2022.855091
  3. Pias, S. B. H., Huang, R., Williamson, D., Kim, M., & Kapadia, A. (2024). The impact of perceived tone, age, and gender on voice assistant persuasiveness in the context of product recommendations. In Proceedings of the 6th ACM Conference on Conversational User Interfaces (CUI ’24). ACM. https://doi.org/10.1145/3640794.3665545
  4. Tolmeijer, S., Zierau, N., Janson, A., Wahdatehagh, J. S., Leimeister, J. M., & Bernstein, A. (2021). Female by default? Exploring the effect of voice assistant gender and pitch on trait and trust attribution. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems (CHI EA ’21). ACM. https://doi.org/10.1145/3411763.3451623
  5. West, M., Kraut, R., & Chew, H. E. (2019). I’d blush if I could: Closing gender divides in digital skills through education. UNESCO & EQUALS Skills Coalition.
  6. Zhang, Q., Yang, X. J., & Robert, L. P., Jr. (2025). Artificial intelligence voice gender, gender role congruity, and trust in automated vehicles. Scientific Reports, 15, Article 16364. https://doi.org/10.1038/s41598-025-00884-9
Disclaimer

This is a preliminary research note, not a peer-reviewed study. It reports observations from a single deployment, and its estimates carry the uncertainty described in the Scope section. The results are specific to the portfolios, scripts, channels, and period studied; they are not a prediction of what any other deployment will show, and actual results will vary by portfolio, configuration, and context. Nothing here is a guarantee of performance, and the analysis may be revised as more data is collected. All data underlying it is aggregated and anonymised, identifies no individual or client, and is used with the clients’ agreement.

This document is provided for informational purposes only and does not constitute legal or regulatory advice. Motivated by Krew’s Collaborative Assurance Framework’s principles of shared accountability and streamlined assurance, we partner with our customers, combining our AI-driven credit servicing platform’s built-in security safeguards with each customer’s own controls, system configurations, and user training, to help manage regulatory risk. It remains the responsibility of each customer to ensure compliance with all applicable federal, state, and local statutes. All rights, responsibilities, and liabilities of Krew in relation to its customers are governed exclusively by the terms of Krew’s customer agreements. This document neither forms part of, nor alters, any contractual agreement between Krew and its customers.