When Is the Best Time to Call?

Every outbound operation has a folk theory of when to call, and it is rarely tested. This preliminary research note builds a model that ranks the coming week’s calling times for each consumer, tests it on the calls made after it was built, and finds that the right time to call depends on the person, not the hour.

Preliminary Research Note · September 2026

There is no good hour to call. There is a good hour to call a particular person.

A dial placed in the slot the model ranked first reached a person in a real conversation more often than the same consumer’s dials in other slots. The industry’s best-hour lookup table, on the same consumers, did no better than chance.

The question

Every outbound operation has a folk theory of when to call, whether the call is an origination follow-up, a servicing touch or a collections attempt: evenings and weekends, or late mornings, or whenever the last conversation happened. The theory varies by shop and is rarely tested. The question that matters operationally is narrower than “when is the best time to call,” because attempts are limited everywhere: by Regulation F’s call-frequency presumption for collections calls, and by each team’s own cadence rules for servicing and originations. The question is which of a handful of slots, for which consumer, in which order.

This note describes a per-consumer call-timing model built by Krew on the voice agent’s outbound dials, and its validation on production traffic. It is a longitudinal study, meaning it follows accounts over time, of historical engagements run by financial services firms on Krew’s platform, which places outbound calls for originations, servicing and collections. Only the customers who opted into the study are included. The historical call times were set by each participating firm’s own scheduling, not by the model, and the model was trained on a subset of those dials and tested on the dials that came after.

The dial timeline was captured and scored with Observer, which records every call, text and email and grades each one against the same scorecards. The model was built and validated with Progenitor, which reports what is actually moving outcomes and turns the result into workflow rules. Every figure below comes from that one timeline.

Results are given as odds ratios. An odds ratio of 1.37 means a call in the model’s preferred slot had 37% higher odds of turning into a conversation than the same person’s calls at other times, and 1.00 means no difference. The 95% CI is the range the true value is likely to fall in.

What we found

First, the model’s preferred slot reaches people. On production dials, a dial placed in the slot the model ranked first had 37% higher odds of reaching a person in a real conversation than the same consumer’s dials in other slots (odds ratio 1.37, 95% CI [1.23, 1.53]). Each consumer is compared only with themselves, so the model gets no credit for knowing which people are easy to reach, only for knowing when to reach them. Repeated over four back-to-back periods, each scored by a model that had seen only the data before it, the combined odds ratio is 1.21 (95% CI [1.08, 1.35], p = 0.0006), and the periods agree with one another. A fifth to a third more conversations from the same dial is the range the evidence supports.

Second, the industry’s standard approach is no better than chance. The usual rule, a lookup table of the best hour and weekday for each firm and timezone, ranked the same consumers’ slots with an odds ratio of 1.00, a coin flip. The gain comes from knowing the person: who is being called and how earlier attempts went carry most of the signal, and timing adds a steady increment on top. There is no good hour to call; there is a good hour to call a particular person.

Third, the output is a weekly plan built inside the rules, not a score. For every open account the firm may lawfully call, the model produces a ranked plan of slots for the week, capped at the firm’s own attempt limit for that use case and inside Regulation F’s call-frequency presumption where it applies, permitted calling hours and any times the consumer has said are inconvenient, spread over four to six distinct days with at most two attempts per day and same-day attempts at least three hours apart. The plan is used in rank order and stops at the first conversation. It uses no protected characteristics, and its result does not depend on which firm works the account.

The model sits inside the workflow rather than beside it: its inputs are the platform’s own touch timeline, and its output is a schedule the voice agent executes.

What the research says

Household history beats the population rule, and every attempt has to count

  1. 1. Population-level calling rules are old, and they do not travel

    Weeks, Kulka and Pierson (1987) established the classic result on a national telephone survey: the probability of an answer and a completed interview on the first attempt was substantially higher on weekday evenings and weekends than during weekday daytime, and concentrating calls there reduced callbacks. Shino and McCarty (2020), on eight years of call records from a modern survey, found the opposite: afternoon shifts were as productive as evenings and early weekdays beat weekends.

    The heuristic is unstable across populations and eras, which is what a firm-by-timezone lookup table encodes and what its chance-level result below reflects.

  2. 2. Household-level history beats the population rule

    Wagner (2013) fit models estimating each household’s probability of contact in each of four calling windows and prioritized cases in real time by their highest predicted window. In the telephone arm of that study the adaptive protocol raised the contact rate by about a fifth relative to the fixed protocol and cut calls per interview by more than a tenth. Lipps (2012), on Swiss panel call records, found that choosing the first call time from the household’s own prior contact history produced earlier successful contact at lower fieldwork cost. Durrant, Maslovskaya and Smith (2017) found that the most recent call outcome carried more predictive weight than either geography or older history. Vicente, Marques and Reis (2017), on mobile numbers, found that day of week, the gap between attempts and the sequence of prior outcomes all carried signal.

    Each of these is a feature the model below engineers from the platform’s own dial log.

  3. 3. Every attempt has to count

    Triplett (2002) documented across telephone surveys that first attempts yield about a fifth of eventual completions and that the return from each further attempt falls steadily. For collections calls, Regulation F (12 CFR § 1006.14(b)) sets a call-frequency presumption that every plan in this note sits inside, alongside each firm’s own limits for servicing and originations and any times a consumer has said are inconvenient.

    With few attempts available, which slot each one lands in decides how many of them reach a person.

Literaturefour findings

Household history beats the population rule, and every attempt has to count

The published results this note compares its own findings with.

+⅕

Prioritizing each household by its best predicted calling window raised the contact rate by about a fifth and cut calls per interview by more than a tenth.

Wagner (2013)

recent

The most recent call outcome carried more predictive weight than geography or older history.

Durrant, Maslovskaya and Smith (2017)

≈ ⅕

First attempts yield about a fifth of eventual completions, and the marginal yield of each further attempt falls steadily.

Triplett (2002)

7 in 7

Regulation F sets a call-frequency presumption for collections calls. Every plan in this note sits inside it, and inside the firm’s own limits.

12 CFR § 1006.14(b)

The model and its inputs

A contact-timing tool that decides when an account is called, never whether

The model predicts, for each consumer and each candidate slot in the coming week, the probability that a dial in that slot produces an engaged conversation. A dial succeeds only when the consumer came on the line and said at least two things; a mailbox greeting, a call-screening assistant, a silent or machine-only line, and a pickup that ends after a single word all count as failed attempts at that time, and the voice agent discloses nothing about the account until the consumer has been verified. Whether a dial succeeded is judged from the consumer’s own words in the transcript, checked against the platform’s voicemail and call-screening tags.

The production model combines several simpler models that share the same inputs. It was chosen from a broad search, and the simpler candidates held up best, because the timing signal is subtle next to who is being called. Every version is saved with its inputs and its test scores.

The model draws on five groups of features: when the candidate slot falls, in the consumer’s local time; how the consumer’s earlier attempts went in and around that slot; how they went overall, and how recently; recent text and email activity on the account; and coarse, non-demographic account attributes.

Excluded by design: date of birth and age, name, city, postal code, state beyond its timezone mapping, region, language, personal situation, employer, and any consumer-report data. The model is a contact-timing tool; it does not score creditworthiness, and it is not used to decide whether an account is worked, only when.

The model and its inputsfive feature groups

What the model draws on, and what it is built without

A contact-timing tool: it decides when an account is called, never whether, and scores no creditworthiness.

Slot

When the candidate slot falls, in the consumer’s local time

Slot-specific history

How the consumer’s earlier attempts went in and around that slot

Attempt history

How the consumer’s earlier attempts went overall, and how recently

Other channels

Recent text and email activity on the account

Account context

Coarse, non-demographic account attributes

Excluded by design

Date of birth and ageNameCity, postal code, regionState beyond its timezoneLanguagePersonal situationEmployerAny consumer-report data
Validation on production traffic

The model’s preferred slot reaches more people than the same consumer’s other slots

The model was trained on dials from June and July 2026 and validated on production dials from 1 August to 11 September 2026. Two results matter. The first is whether the model’s scores tell apart the dials that reached a person from the dials that did not. The second, and the one a decision rests on, is the slot lift within each consumer: for consumers who were dialed at two or more different times during the window, the model ranked those times using only history from before the window, and their dials in its preferred slot are compared with their own dials at other times. Because each person is compared only with themselves, the model cannot earn credit for knowing who is reachable; it can only earn credit for knowing when.

The model’s preferred slot carries an odds ratio of 1.37 (95% CI [1.23, 1.53]) against the consumer’s other tried slots, p < 0.0001. A history-aware single model’s preferred slot carries 1.29 (95% CI [1.16, 1.44]). The lookup table’s preferred slot, by firm, timezone, hour and weekday, carries 1.00 (95% CI [0.90, 1.12], p = 0.97). The first dial of the window against later dials carries 0.66 (95% CI [0.59, 0.74]).

Attempt order works against the model, not for it. First dials in the window reach a person less often than later ones, so the lift is not an artefact of the model preferring early attempts. Asked which of two of a consumer’s slots produced the conversation, the model gets it right 1.15 times as often as chance (95% CI [1.10, 1.20]).

The lookup table is the control. It is the conventional population-level rule, and on the same consumers it ranks slots no better than a coin. The whole of the lift is personalization.

Validationproduction dials, 1 Aug to 11 Sep 2026

The model’s preferred slot reaches more people than the same consumer’s other slots

Odds of a real conversation in the model’s preferred slot against the same consumer’s other tried slots, 95% CI. Each consumer is compared only with themselves. 1× means no difference.

Model’s preferred slot

1.37×

95% CI [1.23, 1.53], p < 0.0001

Population lookup table

1.00×

95% CI [0.90, 1.12], p = 0.97

Confirmation across rolling windows

The lift repeats in every production window

A single period can flatter a model that happens to suit the season. So the model was rebuilt and re-scored over four back-to-back periods from 1 July to 11 September, each time using only the data before that period. The lift repeats in every period large enough to measure it, the periods agree with one another (a test for differences between them gives p = 0.36), and combined across all four the odds of a conversation in the preferred slot are a fifth higher: 1.21 (95% CI [1.08, 1.35], p = 0.0006).

The three periods large enough to report alone carry odds ratios of 1.22 (95% CI [1.01, 1.47]) for 15 to 31 July, 1.28 (95% CI [1.04, 1.57]) for 1 to 15 August, and 1.20 (95% CI [1.00, 1.44]) for 16 August to 11 September. The first period, 1 to 14 July, was scored by a model trained on a single month and is too small to be informative on its own; it is included in the combined estimate.

The two headline figures fit together. The main period, 1 August to 11 September, trained through July, gives an odds ratio of 1.37; the rolling periods, with shorter training histories and shorter scoring periods, give 1.20 to 1.28 each and 1.21 combined. The ranges overlap, and the two designs answer slightly different questions: the main period is the best single estimate of the model as deployed, and the rolling design is the cautious estimate of how the lift holds up as the book changes.

A fifth to a third more conversations from the same dial is the range this note stands behind. Combining several models is what makes the lift hold up across periods, which is why the production model is the combination rather than the best single model.

Rolling validation1 Jul to 11 Sep 2026

The lift repeats in every production window

Each period scored by a model that had seen only the dials before it. Odds ratio, each consumer compared with themselves, 95% CI.

Where the signal comes from

Most of the predictable part of engagement is who is being called and how earlier attempts went

How well each model tells reachable slots apart is scored on a scale where a coin flip is zero (the Gini index), and as the improvement over simply predicting the average rate (log loss). Both tell the same story: a population-level lookup table is barely above a coin flip, at 0.02, and slightly worse than the average rate; timing and account context without attempt history is a modest improvement, at 0.10; a model that knows the consumer’s attempt history tells the slots apart, at 0.30; and the production model adds a little more, at 0.32 and a 4.0% improvement over the average rate.

Timing adds a small, consistent amount on top of history, which is the shape the survey research predicts: Durrant et al. (2017) found the most recent call outcome carried the weight, and Wagner (2013) found the gains came from household-level prediction rather than from a better population rule.

A timing-only model is close to chance, and that is the point. The lift is not in the hour; it is in the hour for this consumer, given what the platform already knows about them.

Where the signal comes fromproduction dials, 1 Aug to 11 Sep 2026

A history-aware model separates reachable slots; a population-level rule does not

How well each model tells reachable slots apart, on a scale where a coin flip is zero (the Gini index), and the improvement over predicting the average rate (log loss).

What the model produces

A ranked week for every account, built inside the rules

For every open account with at least one outbound dial that the firm may lawfully call, the model produces a ranked plan of slots for the week, capped at the firm’s own attempt limit, with a predicted conversation probability per slot: at most two per day, same-day attempts at least three hours apart, spanning four to six distinct days, all within permitted calling hours in the consumer’s local time and outside any time the consumer has said is inconvenient. The plan is used in rank order, stops at the first conversation, and is suspended by a cease request, a dispute or a do-not-call instruction.

The top-ranked slot is shared by many consumers, because the strongest population-level slot is a reasonable first guess for a consumer with little history. The later ranks are where the personalization lives: they vary with the consumer’s own pickups, conversations, voicemails and no-answers, the hour and weekday of the last engagement, and what the other channels have done in the past week. Only slots with enough dial support to be estimated reliably are recommendable.

The model does not depend on knowing the firm. Rebuilding the model without the firm as an input, on the same period, leaves every result unchanged: the same scores, the same odds ratio of 1.37, the same accuracy on pairs of slots. A standard test of how much the model leans on the firm input returns zero. The account context and attempt-history features already carry what the firm label would add, so the result does not depend on which firm works the account.

The scores are used for ordering. How reachable a book is drifts over a season as it ages, and the model is re-scored weekly to keep up. Its ranking of slots stays stable across the rolling periods, and the scores are used to order slots and to compare consumers, not as a forecast of how many calls will connect.

What the model producesa ranked week per account

The top-ranked slot is shared; the later ranks are personal

Share of scored consumers whose first-ranked slot falls on each local-time slot. Later ranks vary with the consumer’s own pickups, conversations, voicemails and no-answers. Only slots with enough dial support are recommendable.

Implications for deployment

The result is a schedule, and the design choices follow from how it was validated

The plan slots into the attempt limits a dialer already respects. What changes is which slot each attempt lands in, and that only works when calls, texts and emails write to the one timeline the model reads.

On Krew, Progenitor re-scores every open account each week under the firm’s policy and hands the ranked plan to the workflow layer. Observer scores every resulting call, so the plan is measured on the book it actually runs on.

  1. 01

    Replace the lookup table, not the dialer

    A population-level best-hour rule is no better than chance on the same consumers. The model’s plan slots into the attempt limits the dialer already respects; what changes is which slot each attempt lands in.

  2. 02

    Spend the first attempt on the first-ranked slot

    The headline lift was measured on the top-ranked slot. A workflow that dials rank one first, and stops at the first conversation, captures the value; a workflow that dials the slots in calendar order captures less.

  3. 03

    Feed the model the whole timeline

    The signal is in attempt history and in the other channels: recent text and email activity informs the ranking, and a conversation on any channel is the event that stops the plan. The model performs as validated only when the voice agent, texts and emails write to one timeline it can read.

  4. 04

    Re-score weekly, retrain on a rolling window

    The rolling validation is the deployment pattern: train on everything before the week, score the week, repeat. The lift held across four consecutive windows with that discipline.

  5. 05

    Place fewer attempts better, not more of them

    Four to six distinct days, at most two attempts a day, three hours apart: the plan is spaced the way the contact literature finds productive, and it sits inside the firm’s own attempt limit for the use case, Regulation F’s frequency presumption where it applies, permitted hours and any times the consumer has said are inconvenient. A firm with a lower attempt limit keeps the ranked top and drops the rest.

  6. 06

    Confirm it on your own book

    Comparing each consumer with themselves removes the effect of who gets called. A controlled comparison of model-timed against scheduler-timed dials, within the firm’s normal attempt limits, is the clean measurement of the remaining question, and the platform’s workflow layer supports it.

Scope and interpretation

What the validation covers, and how to read the numbers

This is a validation over time on the customers who opted into the study. Their accounts are followed over time, and the results describe that group rather than every firm on the platform. Historical call times were chosen by each participating firm’s scheduler, and the model was trained on a subset of those dials and validated on the dials that came after it.

Comparing each consumer with themselves is the test a decision rests on. It removes the effect of who gets called, so the model cannot earn credit for knowing which consumers are reachable, and the order of attempts runs against it. A controlled rollout is the recommended next step for a firm that wants to measure what happens when the model, rather than the scheduler, sets the times.

The lift was measured on the top-ranked slot. The plan is used in rank order for that reason, and the two headline figures, an odds ratio of 1.37 on the main period and 1.21 combined across the rolling periods, mark out the range rather than contradict it.

Support. The platform dials only within permitted hours in the consumer’s local time, and the model scores only slots that have been tried often enough to estimate reliably. Success requires two consumer turns, which guards against a greeting or a screening assistant being counted as a conversation.

Source. Krew production database, outbound voice calls with a true dial time, 1 April to 11 September 2026. Dial time is the platform’s scheduled time, which matches the carrier audit trail to within a minute from April onward. Internal test accounts are excluded. The study population is the subset of customers who opted into it. The data is aggregated and anonymised, identifies no individual consumer or firm, and is used with the participating firms’ agreement.

Training and validation. Training subset: 1 June to 31 July 2026, a little over half of the dials. Main validation window: 1 August to 11 September 2026. Rolling confirmation: four windows (1 to 14 July, 15 to 31 July, 1 to 15 August, 16 August to 11 September), each trained only on data before it. April and May are used as attempt history rather than as training labels, because that book differed materially from the later period.

Definitions and estimation
  • Success means an engaged conversation: the consumer came on the line and produced two or more turns and five or more words. Mailboxes, call-screening assistants, silent or machine-only lines, and single-turn pickups are failures at that time. Detection runs on the consumer’s transcript lines with the platform’s voicemail and screening flow tags and the LLM contact labels applied on top.
  • Model. A proprietary ensemble of several learners on a shared feature set, selected from a broad model search on the same production window.
  • Separation is AUC and log loss on the production window, reported here as Gini and as improvement over the constant rate.
  • The within-consumer slot lift is a Mantel-Haenszel odds ratio across consumers dialed in two or more distinct slots, with 95% CIs; rolling windows are pooled by Mantel-Haenszel with a heterogeneity test. Concordance is the share of a consumer’s (conversation slot, other slot) pairs ordered correctly, reported relative to chance.
  • Plan. A ranked list of slots per open account, capped at the firm’s attempt limit: at most two per day, same-day attempts at least three hours apart, four to six distinct days, restricted to slots with enough dial support to be estimated reliably. Used in rank order; stops at the first conversation.
Design parametersKrew production data

Trained on the dials before, validated on the dials after

Outbound voice calls with a true dial time, 1 April to 11 September 2026. April and May serve as attempt history, not as training labels.

Prepared by

Michael Goh

Michael Goh

Cofounder, Member of Technical Staff

Formerly led speech benchmarking at Artificial Analysis. Previously a FinTech VC and Bain consultant. MS Computer Science, University of Chicago; BA Economics, University of Oxford.

Rafael Khaykin

Rafael Khaykin

Cofounder, Member of Technical Staff

Formerly a research engineer at Figure Robotics, working on dexterity algorithms and latency research. Computer Science, University of British Columbia.

Alice Zhang

Alice Zhang

Head of Product Compliance, Member of Technical Staff

Formerly advised the United Nations on AI Ethics and Safety for defense industries. Previously In-House Counsel at a FinTech VC. BA (LLB) Jurisprudence, University of Oxford.

References
  1. Durrant, G. B., Maslovskaya, O., & Smith, P. W. F. (2017). Using prior wave information and paradata: Can they help to predict response outcomes and call sequence length in a longitudinal study? Journal of Official Statistics, 33(3), 801–833. https://doi.org/10.1515/jos-2017-0037
  2. Consumer Financial Protection Bureau (2021). Regulation F, telephone call frequency limits, 12 C.F.R. § 1006.14. https://www.ecfr.gov/current/title-12/chapter-X/part-1006/subpart-B/section-1006.14
  3. Lipps, O. (2012). A note on improving contact times in panel surveys. Field Methods, 24(1), 95–111. https://doi.org/10.1177/1525822X11417966
  4. Shino, E., & McCarty, C. (2020). Telephone survey calling patterns, productivity, survey responses, and their effect on measuring public opinion. Field Methods, 32(3), 291–308. https://doi.org/10.1177/1525822X20908281
  5. Triplett, T. (2002). What is gained from additional call attempts and refusal conversion and what are the cost implications? Urban Institute working paper.
  6. Vicente, P., Marques, C., & Reis, E. (2017). Effects of call patterns on the likelihood of contact and of interview in mobile CATI surveys. Survey Methods: Insights from the Field. https://doi.org/10.13094/SMIF-2017-00003
  7. Wagner, J. (2013). Adaptive contact strategies in telephone and face-to-face surveys. Survey Research Methods, 7(1), 45–55.
  8. Weeks, M. F., Kulka, R. A., & Pierson, S. A. (1987). Optimal call scheduling for a telephone survey. Public Opinion Quarterly, 51(4), 540–549. https://doi.org/10.1086/269056
Disclaimer

This is a preliminary research note, not a peer-reviewed study. It reports a validation on historical engagements from an opt-in subset of the platform, and its estimates carry the uncertainty described in the Scope section. Historical call times were set by each participating firm’s own scheduling; no dial time was assigned by the model, and a controlled rollout remains the recommended next step. The results are specific to the firms, portfolios, workflows, channels, and period studied; they are not a prediction of what any other deployment will show, and actual results will vary by portfolio, configuration, and context. Nothing here is a guarantee of performance, and the analysis may be revised as more data is collected. All data underlying it is aggregated and anonymised, identifies no individual or firm, and is used with the participating firms’ agreement.

This document is provided for informational purposes only and does not constitute legal or regulatory advice. Motivated by Krew’s Collaborative Assurance Framework’s principles of shared accountability and streamlined assurance, we partner with our customers, combining our AI-driven credit servicing platform’s built-in security safeguards with each customer’s own controls, system configurations, and user training, to help manage regulatory risk. It remains the responsibility of each customer to ensure compliance with all applicable federal, state, and local statutes. All rights, responsibilities, and liabilities of Krew in relation to its customers are governed exclusively by the terms of Krew’s customer agreements. This document neither forms part of, nor alters, any contractual agreement between Krew and its customers.