Whom Should You Dial?
Every outbound operation has more accounts to reach than attempts to reach them with, and the usual answer is a queue. This preliminary research note builds a score of how likely the next dial is to reach the consumer, tests it on the calls made after it was built, and finds that the consumer’s own attempt history tells the workflow whom to call next.
Attempts are a budget. The value of the week depends on which accounts receive them.
The score ranks accounts by how likely the next call is to reach the consumer. The top-scoring accounts engaged at six times the rate of the bottom-scoring accounts. It learns from how each consumer’s earlier calls went, not from which firm is calling, so it works on any firm’s book without retraining.
Every outbound operation, whether it is following up an application, servicing an account or working a balance, has more accounts to reach than attempts to reach them with. The attempts are limited, by rule in collections and by cost everywhere, and the question a workflow answers each week is which accounts should receive them. The conventional answer is a queue, oldest first or largest first. This note describes a better one.
Krew built a per-attempt model of the probability that an outbound dial from the voice agent ends in an engaged conversation, and validated it on production traffic. It is the companion to the best-time-to-call model: that model answers when to dial a consumer, this one answers whom to dial, and the two are built on the same data, the same definition of success and the same exclusions. It is a longitudinal study, meaning it follows accounts over time, of historical engagements run by financial services firms on Krew’s platform, which places outbound calls for originations, servicing and collections. Only the customers who opted into the study are included. The historical choice of whom to dial and how often was made by each participating firm’s own scheduling, not by the model, and the model was trained on a subset of those dials and tested on the dials that came after.
The dial timeline was captured and scored with Observer, which records every call, text and email and grades each one against the same scorecards. The model was built and validated with Progenitor, which reports what is actually moving outcomes and turns the result into workflow rules. Every figure below comes from that one timeline.
Results are given relative to the overall rate. A decile at 2.38× reached a person in an engaged conversation at 2.38 times the rate of all dials in the window, and 1.00× means no different from the average.
First, the score separates reachable attempts from unreachable ones. On production dials, the top-scoring tenth of dials reached a person in an engaged conversation at 2.4 times the overall rate and the bottom tenth at 0.4 times, a six-fold spread. The top fifth of dials by score produced two in five of all engaged conversations and a similar share of all promises to pay. The separation held in each of four consecutive production periods, each scored by a model trained only on the dials before it.
Second, the signal is the consumer’s own attempt history, which is what makes the score portable and auditable. How earlier attempts went, how recently, and what the last one produced carry the large majority of the model’s signal. Features that would let the model learn a particular firm’s dialing strategy, and every timing feature the best-time model owns, are excluded at no cost to accuracy, so the score transfers to a new firm or a new book without retraining and cannot encode one firm’s habits as another’s plan. It uses no protected characteristics, no proxies for them, and no consumer-report data.
Third, the score is a reachability tool that lives in the workflow. It ranks accounts by the likelihood that the next attempt reaches the consumer; it does not score creditworthiness, it is never used to decide whether an account is eligible for anything, and it never withholds a communication a consumer is entitled to. Used with the best-time model, it decides which accounts receive attempts this week and the best-time model decides in which slots, both inside the firm’s attempt limits and the consumer’s stated preferences.
The score is re-computed after every attempt, because each attempt is new evidence about the consumer. That is the design this note describes.
Prioritizing by predicted propensity is established practice, and the attempt budget is now a rule
1. Responsive design is established practice
Groves and Heeringa (2006) formalized responsive survey design, in which the probability of a successful contact is modeled for each case from the record of earlier attempts and the effort of the field period is redirected toward the cases where it will change the outcome. Wagner (2013) implemented the idea for telephone contact, estimating each household’s contact probability and prioritizing cases accordingly, and raised the contact rate by about a fifth while cutting calls per interview by more than a tenth.
A propensity-to-engage score is the same idea applied to a firm’s weekly attempt budget.
2. The most recent attempt is the most informative feature
Durrant, Maslovskaya and Smith (2017) found, on longitudinal call records, that models built only on design and geographic covariates predicted contact outcomes poorly, that prior history added modest lift, and that conditioning on the most recent call outcome gave the clearest improvement. Lipps (2012) found that a household’s own prior contact pattern predicted earlier successful contact at lower cost. Vicente, Marques and Reis (2017) found that a run of voicemail or no-answer outcomes sharply lowered the probability of a later interview.
These are the features the model below engineers from the platform’s dial log, and the reason it is re-scored after every attempt.
3. Attempts are a budget, and the budget is now a rule
Triplett (2002) documented across telephone surveys that first attempts yield about a fifth of eventual completions and that each further attempt yields less. For collections calls, Regulation F (12 CFR § 1006.14(b)) turned that efficiency argument into a call-frequency presumption, and servicing and originations teams run their own cadences.
When the number of attempts is fixed, the value of the week depends on which accounts receive them.
Prioritizing by predicted propensity is established practice, and the budget is now a rule
The published results this note compares its own findings with.
Estimating each household’s contact probability and prioritizing cases by it raised the contact rate by about a fifth and cut calls per interview by more than a tenth.
Wagner (2013), after Groves and Heeringa (2006)
Models built on design and geography alone predicted contact poorly; conditioning on the most recent call outcome gave the clearest improvement.
Durrant, Maslovskaya and Smith (2017)
First attempts yield about a fifth of eventual completions, and Regulation F turns that efficiency argument into a call-frequency presumption for collections calls.
Triplett (2002); 12 CFR § 1006.14(b)
A contact-prioritization tool that decides which accounts the next attempts reach, never whether
The model predicts, for each account and its next outbound dial, the probability that the dial produces an engaged conversation. Success is defined exactly as in the best-time model: the consumer came on the line and said at least two things. A mailbox greeting, a call-screening assistant, a silent or machine-only line and a pickup that ends after a single word are all failures at that attempt, not successes and not dropped rows, and the voice agent discloses nothing about the account until the consumer has been verified.
The score answers whom to dial, not when. It deliberately excludes every timing feature, which the best-time model owns; the firm, the timezone and recent text and email activity, so that the model cannot learn one firm’s dialing strategy and carry it to another; and every protected characteristic and proxy. What remains is a compact feature set in two groups: how the consumer’s earlier attempts went, how recently, and what the last one produced; and coarse, non-demographic account attributes.
The production model combines several simpler models trained on that feature set, chosen from a broad search on the same production period. The simpler, smoother candidates held up best, and the combination led the best single model in every rolling period. Every version is saved with its inputs and its test scores.
Excluded by design: date of birth and age, name and any inferred gender, city, postal code, state, region, language, personal situation, employer, any consumer-report data, the firm working the account, the timezone, and recent text and email activity. The score is a contact-prioritization tool: it does not score creditworthiness, it is not used to decide eligibility, pricing or terms, and it never withholds a communication a consumer is entitled to receive.
The score reads the consumer’s own attempt history
One account’s dials over four weeks, and the score after each one. Unanswered attempts lower it, a conversation raises it, and the levels here are illustrative.
How many were picked up
and how many went unanswered, to voicemail or to a screening assistant
What the last attempt produced
the single most informative feature, which is why the score is refreshed after every dial
How long since the consumer last engaged
a conversation weeks ago counts for less than one last week
Coarse account context
non-demographic attributes of the account, adding a small remainder
Never read
The top decile of dials engages at 2.4 times the overall rate; the bottom decile at 0.4 times
The model was trained on dials from June and July 2026 and validated on production dials from 1 August to 11 September 2026, each dial scored with only the history available before it. Two results matter. The first is separation: whether the score distinguishes dials that reached a person from dials that did not. The second, and the operational one, is what happens when dials are worked in score order: how many of the conversations, and how many of the promises to pay, fall in the top of the ranking.
Sorted into ten equal groups by score, the top group reached a person at 2.38 times the overall rate and the bottom group at 0.39 times, with the groups between them falling in order: 1.50, 1.19, 1.07, 0.90, 0.83, 0.67, 0.56 and 0.50 times. How well the score tells the two kinds of dial apart, on a scale where a coin flip is zero (the Gini index), is 0.30 on the main period and 0.29 averaged across the four rolling periods.
The separation repeats. Across four consecutive production periods from 1 July to 11 September, each scored by a model trained only on the dials before it, the separation above chance held steady and the combined model led the best single model in every period.
A six-fold spread between the most and least reachable tenth of dials, from attempt history and account context alone.
The top decile engages at 2.4 times the overall rate; the bottom decile at 0.4 times
Engaged-conversation rate in each tenth of dials by score, relative to the overall rate. Every dial scored with only the history available before it.
Top decile
2.38×
the overall rate
Bottom decile
0.39×
a six-fold spread
Separation above chance
0.30
Gini; 0.29 across rolling windows
Two dials in five that reach a person come from the top fifth of scores
The ranking concentrates the outcomes that matter. The top fifth of dials by score produced 39% of all engaged conversations and 38% of all promises to pay in the period; the top three deciles produced half of the conversations, and the bottom three deciles about one in seven.
An attempt budget that covers a fifth of the candidate dials reaches roughly twice as many people when spent from the top of the ranking as when spent at random. The dashed line on the chart is the unranked queue: every tenth of dials worked captures a tenth of the conversations. The curve above it is the same dials worked in score order.
Promises to pay are downstream of the conversation. The score does not predict them; it predicts the conversation, and the promises follow because the conversations do. They are reported here as where they fall in the ranking, not as what the model forecasts.
Probabilities are used for ordering. How reachable a book is drifts over a season as it ages, and the model is re-scored after every attempt and refit on a rolling window to track it. The ranking is stable across periods, and the scores are used to order accounts and compare them rather than as a forecast of the conversation rate.
Two dials in five that reach a person come from the top fifth of scores
Cumulative share of engaged conversations captured as dials are worked in score order. The dashed line is what an unranked queue captures.
The large majority of the signal is attempt history, and that has three consequences
The large majority of the model’s signal is attempt history: how many earlier attempts were picked up, how many went unanswered, what the most recent attempt produced, and how long ago the consumer last engaged. Account context adds a small remainder. That shape is exactly what the survey-methodology literature predicts, and it has three consequences for deployment.
The score is portable. Adding the firm, the timezone or the other channels’ recent activity to the feature set changed separation by nothing measurable, so excluding them costs nothing and buys a score that transfers to a new firm or a new book without retraining and cannot encode one firm’s dialing habits as another firm’s plan.
The score is robust to the definition of success. Retraining on a stricter target, conversations with a verified consumer only, ranks the same accounts at the top: the top decile under either definition is the same population. The engaged-conversation target is kept because it is the outcome the voice agent records at the moment it happens.
The score is worth refreshing. Because the signal is the most recent attempts, a score taken once at the start of a period is weaker than a score refreshed after every dial, though even the single early score separates the book: the top decile of accounts scored once, at their first dial, went on to a conversation at 2.2 times the rate of all accounts and the bottom decile at 0.4 times, and the top two deciles of accounts held a third of all conversing accounts and 38% of promises.
Rescoring after every attempt is the intended use.
A single score per account, taken at its first dial, still separates the book
Share of accounts reaching a conversation in the window, relative to all accounts. Rescoring after every attempt is the intended use and is stronger.
The score is a ranking, and the design choices follow from how it was validated
The score orders attempts within the limits a dialer already respects. What changes is which accounts the attempts reach, and that only works when calls, texts and emails write to the one timeline the model reads.
On Krew, Progenitor re-scores every open account after each attempt under the firm’s policy and hands the ranking to the workflow layer. Observer scores every resulting call, so the ranking is measured on the book it actually runs on.
- 01
Spend the attempt budget from the top of the ranking
Two in five conversations and a similar share of promises to pay come from the top fifth of dials. A budget that covers a fifth of the candidate attempts reaches roughly twice as many people when spent in score order as when spent in queue order.
- 02
Pair whom with when
The propensity score decides which accounts receive attempts this week; the best-time model decides in which slots. The two are built on the same timeline and the same definition of success, exclude the same characteristics, and divide the features between them so that neither learns what the other owns. Together they turn a weekly attempt budget into a ranked, timed plan.
- 03
Re-score after every attempt
Each dial is evidence. The score that ranked an account first on Monday is stale by Tuesday if Monday’s dial went unanswered. A workflow that re-scores on every outcome captures the lift; a workflow that scores once a week captures less.
- 04
Keep the exclusions
The score uses no protected characteristics, no proxies and no consumer-report data, and the firm, timezone and channel exclusions keep it from learning one firm’s dialing habits. It is a contact-prioritization tool, and it stays one: it is not used to decide eligibility, pricing or terms.
- 05
Use the score to prioritize, never to withhold
The score orders attempts within a budget. It does not decide eligibility, pricing or terms, and it never suppresses a communication a consumer is entitled to receive. Attempt limits, cease requests, disputes and consumer-designated inconvenient times sit in the workflow’s scheduler, which both scores feed and neither overrides.
- 06
Confirm it on your own book
The historical choice of whom to dial was made by each firm’s scheduler, so attempt-history features partly reflect its choices. A controlled comparison of score-ranked against scheduler-ranked attempts, within the firm’s normal attempt limits, is the clean measurement of the remaining question, and the platform’s workflow layer supports it.
What the validation covers, and how to read the numbers
This is a validation over time on the customers who opted into the study. Their accounts are followed over time, and the results describe that group rather than every firm on the platform. Historical dialing was chosen by each participating firm’s scheduler, and the model was trained on a subset of those dials and validated on the dials that came after.
This is a reachability score. It ranks the likelihood that the next attempt reaches the consumer, which is the outcome the workflow controls. Promises to pay and payment plans are downstream of the conversation and are reported here as where they fall in the ranking, not as what the model predicts.
Comparisons are observational. The decile analysis describes what the scheduler’s dials went on to produce when ordered by score. A controlled rollout is the recommended next step for a firm that wants to measure what happens when the score, rather than the scheduler, sets the order.
Success is an engaged conversation. The stricter target of conversations with a verified consumer ranks the same accounts at the top, and the voice agent discloses nothing about the account until the consumer has been verified.
Source. Krew production database, outbound voice calls with a true dial time, 1 April to 11 September 2026. Internal test accounts are excluded. The study population is the subset of customers who opted into it. Data, split and definitions match the best-time-to-call model. The data is aggregated and anonymised, identifies no individual consumer or firm, and is used with the participating firms’ agreement.
Training and validation. Training subset: 1 June to 31 July 2026, a little over half of the dials. Main validation window: 1 August to 11 September 2026. Rolling confirmation: four windows from 1 July to 11 September, each trained only on data before it.
- Success means an engaged conversation: the consumer came on the line and produced two or more turns and five or more words. Mailboxes, call-screening assistants, silent or machine-only lines and single-turn pickups are failures at that attempt. Detection runs on the consumer’s transcript lines with the platform’s voicemail and screening flow tags and the LLM contact labels applied on top.
- Model. A proprietary ensemble of several learners on a shared feature set, selected from a broad model search on the same production window. Timing features, the firm, the timezone, recent text and email activity, protected characteristics, proxies and consumer-report data are excluded.
- Separation is AUC and log loss on the production window, reported here as separation above chance (Gini). Decile lift is the engaged rate in each tenth of dials by score relative to the overall rate; cumulative capture is the share of engaged conversations and of promises in the top-scoring share of dials. Account-level results score each account once at its first dial of the window. Feature contribution is permutation importance on the production window.
- Output. A score and a decile for every open account the firm may lawfully call, re-computed after every attempt and refit on a rolling window. The score feeds the scheduler; the scheduler enforces attempt limits, calling hours, consumer-designated inconvenient times, cease requests and disputes.
Trained on the dials before, validated on the dials after
Outbound voice calls with a true dial time, 1 April to 11 September 2026. April and May serve as attempt history, not as training labels.
- Durrant, G. B., Maslovskaya, O., & Smith, P. W. F. (2017). Using prior wave information and paradata: Can they help to predict response outcomes and call sequence length in a longitudinal study? Journal of Official Statistics, 33(3), 801–833. https://doi.org/10.1515/jos-2017-0037
- Groves, R. M., & Heeringa, S. G. (2006). Responsive design for household surveys: Tools for actively controlling survey errors and costs. Journal of the Royal Statistical Society: Series A, 169(3), 439–457. https://doi.org/10.1111/j.1467-985X.2006.00423.x
- Consumer Financial Protection Bureau (2021). Regulation F, telephone call frequency limits, 12 C.F.R. § 1006.14. https://www.ecfr.gov/current/title-12/chapter-X/part-1006/subpart-B/section-1006.14
- Lipps, O. (2012). A note on improving contact times in panel surveys. Field Methods, 24(1), 95–111. https://doi.org/10.1177/1525822X11417966
- Triplett, T. (2002). What is gained from additional call attempts and refusal conversion and what are the cost implications? Urban Institute working paper.
- Vicente, P., Marques, C., & Reis, E. (2017). Effects of call patterns on the likelihood of contact and of interview in mobile CATI surveys. Survey Methods: Insights from the Field. https://doi.org/10.13094/SMIF-2017-00003
- Wagner, J. (2013). Adaptive contact strategies in telephone and face-to-face surveys. Survey Research Methods, 7(1), 45–55.
This is a preliminary research note, not a peer-reviewed study. It reports a validation on historical engagements from an opt-in subset of the platform, and its estimates carry the uncertainty described in the Scope section. The historical choice of whom to dial was made by each participating firm’s own scheduling; no attempt was assigned by the score, and a controlled rollout remains the recommended next step. The results are specific to the firms, portfolios, workflows, channels, and period studied; they are not a prediction of what any other deployment will show, and actual results will vary by portfolio, configuration, and context. Nothing here is a guarantee of performance, and the analysis may be revised as more data is collected. All data underlying it is aggregated and anonymised, identifies no individual or firm, and is used with the participating firms’ agreement.
This document is provided for informational purposes only and does not constitute legal or regulatory advice. Motivated by Krew’s Collaborative Assurance Framework’s principles of shared accountability and streamlined assurance, we partner with our customers, combining our AI-driven credit servicing platform’s built-in security safeguards with each customer’s own controls, system configurations, and user training, to help manage regulatory risk. It remains the responsibility of each customer to ensure compliance with all applicable federal, state, and local statutes. All rights, responsibilities, and liabilities of Krew in relation to its customers are governed exclusively by the terms of Krew’s customer agreements. This document neither forms part of, nor alters, any contractual agreement between Krew and its customers.


