Page contents (Jump to each section)

ClientLab Remodeler Diagnostic

Model methodology and assumption framework

A plain-English account of how we designed the diagnostic, how it works, what it assumes, and why

Prepared by ClientLab Document Version: 1.0 Last Updated: September 2026


Section 1: Purpose of this document

ClientLab built the diagnostic to help established residential remodeling companies assess how effectively they turn incoming demand into profitable, controllable work.

This document explains how we designed the diagnostic.

It covers:

  • what the diagnostic is intended to measure;
  • why the seven diagnostic areas were selected;
  • how the questions were designed;
  • why different types of questions use different response scales;
  • how answers are converted into scores;
  • how each category is normalized;
  • why there is deliberately no single overall score;
  • how individual answers are interpreted beyond their numerical value;
  • how apparently contradictory answers are reconciled;
  • how findings across different areas of the business can be connected without claiming unsupported causation;
  • how published research was used;
  • which parts of the methodology are evidence-backed findings and which are ClientLab modelling decisions;
  • how we tried to reduce commercial and confirmation bias;
  • how the diagnostic was tested;
  • what testing caused us to change;
  • and what the results can and cannot reasonably tell you.

The aim is to make the methodology transparent.

We do not believe that publishing a list of research citations is enough to make a diagnostic trustworthy.

A reader should be able to understand:

What did you measure? Why did you measure it? Why did you ask the question that way? Why does that answer receive that score? How does the score become a conclusion? What evidence supports the conclusion? What assumptions did you have to make along the way?

This document answers those questions.


Section 2: What the diagnostic is and is not

The diagnostic is a structured operational self-assessment for established residential remodeling businesses.

It traces how demand becomes a conversation, how the conversation becomes a qualified opportunity, how opportunities move through sales, and how management tracks their outcomes. It also examines whether the company can select suitable work, the operational effects of those choices, and whether project economics hold up during delivery.

The diagnostic contains:

  • 50 respondent questions in total;
  • 46 scored diagnostic questions;
  • 4 unscored Business Snapshot questions;
  • 7 independently assessed operating areas;
  • answer-level semantic interpretation;
  • within-category pattern logic;
  • cross-category relationship logic;
  • contradiction-reconciliation rules;
  • and deterministic report assembly.

The reporting system produces seven separate category reports, calculates no overall score and does not require a generative AI model to interpret answers.

2.1 What it is

The most technically accurate description is:

A deterministic, criterion-referenced, multi-domain operational self-assessment using ordinal answer-severity coding, independently normalized category exposure indices and rule-based semantic interpretation.

In plain English:

Deterministic

The same set of answers should produce the same underlying findings.

The report is not created by giving the respondent's answers to an AI and asking:

"What do you think is wrong with this company?"

The interpretation rules and approved findings are defined in advance.

Criterion-referenced

The respondent is assessed against defined operating conditions.

The diagnostic does not say:

"You are better than 73% of remodelers."

We do not currently possess a sufficiently representative population dataset to make that statement.

Instead, answers are evaluated against operational criteria such as:

  • how quickly leads receive meaningful engagement;
  • whether qualification is repeatable;
  • whether next actions are secured;
  • whether sales outcomes are visible;
  • whether project standards can be protected;
  • whether normal operations require constant owner intervention;
  • and whether project economics can be compared with what was originally expected.

This is conceptually similar to criterion-referenced assessment: performance is evaluated against pre-defined criteria rather than against the performance of everybody else.

Multi-domain

A remodeling company is not reduced to one number.

A company can simultaneously have:

  • excellent delivery controls;
  • weak lead handling;
  • strong project selection;
  • poor sales visibility;
  • and high owner dependence.

Those are different conditions requiring different interpretations.

Operational

The purpose is not to measure personality, intelligence or general entrepreneurial ability.

It examines observable or reportable operating conditions.

Self-assessment

The diagnostic uses information provided by the respondent.

It is not ordinarily connected directly to:

  • bank accounts;
  • accounting software;
  • CRM history;
  • call logs;
  • project-management records;
  • calendars;
  • job-costing systems;
  • or employee records

to independently verify every answer.

That limitation matters and is discussed later in this document.


Section 3: The core design question

Before writing questions, we had to define what the diagnostic should measure.

The core design objective was:

Assess how an established remodeler handles demand through to profitable work, and identify where the process may waste time or opportunity, narrow project choice, reduce margin or visibility, or limit owner freedom.

This objective covered more than:

"How good is your marketing?"

It also covered more than:

"How good is your CRM?"

It was not designed to answer:

"Which ClientLab features should this company buy?"

The diagnostic needed to remain useful when it identified problems outside ClientLab's scope.


Section 4: How we chose the seven diagnostic areas

We organized the seven areas around a remodeling company's operating journey, independent of a software feature list.

We asked:

What has to work between generating an enquiry and ultimately delivering profitable work through a business that does not depend entirely on its owner?

The result was seven related areas, each measured separately.

The seven-part operating chain

  1. 01Lead Capture & Response
  2. 02Qualification & Opportunity Prioritisation
  3. 03Follow-Up & Sales Progression
  4. 04Visibility & Sales Accountability
  5. 05Project Selectivity & Customer Fit
  6. 06Operational Strain & Owner Dependence
  7. 07Delivery & Margin Reality Check

1. Lead Capture & Response

Core question:

Does incoming demand become a meaningful conversation while the homeowner is still engaged?

This area examines:

  • response speed;
  • response coverage;
  • ability to create two-way conversations;
  • and confidence that paid opportunities are actually being worked.

2. Qualification & Opportunity Prioritisation

Core question:

Can the business identify which opportunities deserve scarce human sales and estimating time before that time is consumed?

This area examines:

  • information available before the first sales conversation;
  • qualification structure;
  • timing of budget discussion;
  • pre-qualification;
  • customer fit;
  • and time consumed by unsuitable enquiries.

3. Follow-Up & Sales Progression

Core question:

Does genuine interest consistently become a defined next step, decision or nurture path?

This area examines:

  • follow-up systems;
  • reminders;
  • ghosting;
  • next-step control;
  • process documentation;
  • and estimate-to-sale conversion.

4. Visibility & Sales Accountability

Core question:

Can management see what is happening well enough to distinguish a bad lead from bad handling, weak qualification, weak follow-up or weak sales execution?

This area examines:

  • source-level conversion;
  • pipeline value;
  • acquisition-cost visibility;
  • salesperson conversion;
  • lost-sale reasons;
  • and outcome recording.

5. Project Selectivity & Customer Fit

Core question:

Does the business have enough commercial choice to protect its standards rather than accepting work because the schedule or cash flow needs feeding?

This area examines:

  • estimate shopping;
  • attitudes around price;
  • willingness to walk away;
  • acceptance of poor-fit work;
  • pricing pressure;
  • customer-fit consequences;
  • recent project quality;
  • pipeline pressure;
  • and uncertainty about future work.

6. Operational Strain & Owner Dependence

Core question:

Can normal work continue through visible systems and defined responsibility, or does the owner remain the person who carries the complete picture?

This area examines:

  • labour availability;
  • labour-driven delays;
  • ability to step away;
  • concentration of knowledge;
  • delegation confidence;
  • interruptions;
  • chasing;
  • handoff quality;
  • and the effect of growth on owner workload.

7. Delivery & Margin Reality Check

Core question:

Does the margin expected when a project is sold survive the reality of delivering it?

This area examines:

  • expected versus actual gross profit;
  • estimated-versus-actual project costing;
  • change-order discipline;
  • delay-related margin loss;
  • and rework or callbacks.

Section 5: Why the seven areas are connected but not combined

The seven categories form an operating chain.

That does not mean they form a simple linear causal chain where Category 1 automatically causes Category 2, Category 2 automatically causes Category 3, and so on.

A remodeling business is more complicated than that.

For example:

  • slow response may contribute to fewer conversations;
  • fewer conversations may reduce the number of good projects available;
  • reduced opportunity may increase pressure to accept weaker-fit work;
  • weaker-fit work may create more operational friction;
  • and operational friction may increase owner involvement.

But those same downstream problems could also be influenced by:

  • market conditions;
  • sales ability;
  • pricing;
  • capacity constraints;
  • management quality;
  • estimating errors;
  • labour shortages;
  • project complexity;
  • customer behaviour;
  • or factors the diagnostic does not measure.

For that reason, the reporting system follows a principle closer to contribution reasoning than exclusive causal attribution.

Contribution analysis is designed around the idea that an intervention or condition can contribute to an outcome without being assumed to be the sole cause of it. Theory-of-change and causal-pathway approaches similarly examine the links between activities, intermediate outcomes and final outcomes while remaining open to alternative explanations and contextual factors.

This principle was explicitly built into the diagnostic.

The system does not claim that weaknesses in the front end of a business are the sole cause of every downstream problem. It identifies where those weaknesses may be contributing to wasted opportunity, weak selectivity, team strain and owner dependence. The Delivery & Margin section separately captures important problems outside ClientLab's direct front-end scope.


Section 6: The "Map the Gap" framework

  1. 01Current state
  2. 02Desired operating state
  3. 03The gap
  4. 04Possible mechanism or consequence
  5. 05Improvement direction
  6. 06ClientLab bridge, where relevant

The reporting methodology uses a structure we refer to internally as Map the Gap.

The framework moves from reported symptoms to operational findings before considering a product recommendation.

The model follows this sequence:

Current state

What does the respondent's own evidence suggest is happening now?

↓

Desired operating state

What would a more controlled version of this area look like?

↓

The gap

What appears to be missing between the two?

↓

Possible mechanism or consequence

How could that gap be contributing to other problems in the business?

↓

Improvement direction

What type of operating capability would reduce the gap?

↓

ClientLab bridge, where relevant

Only after the diagnostic conclusion has been established do we determine whether a ClientLab capability genuinely relates to the identified problem.

The architecture separated the seven category reports and cross-category interpretation from the later ClientLab bridge. That layer covered only problems ClientLab could influence.


Section 7: The questionnaire uses more than one question type

We did not use the same format for all 46 scored questions. That would have been easier, but methodologically weaker.

The questions measure different kinds of information:

  • attitudes;
  • beliefs;
  • confidence;
  • frequency;
  • percentage outcomes;
  • observable process states;
  • binary controls;
  • and behavioural consequences.

Using one answer format for every construct would create artificial uniformity.

Instead, the diagnostic uses several types of response format.


Section 8: Likert-type agreement questions

Some questions use a five-position agreement format such as:

"Every new lead gets an instant response from me or my team, no matter the time of day or night."

The response choices are:

Strongly agree Somewhat agree Neutral Somewhat disagree Strongly disagree

An individual question like this is a Likert-type item, not a full "Likert scale."

A true Likert scale traditionally refers to a composite measure created by combining multiple Likert-type items. Modern measurement literature makes this distinction explicitly.

8.1 Why use agreement questions at all?

Agreement questions are useful when the thing being measured is partly:

  • a belief;
  • a perception;
  • a level of confidence;
  • or a broad operating condition that cannot be reduced naturally to one precise number.

Examples include:

"Every new lead gets an instant response…"

or:

"Most homeowners only care about the lowest price…"

or:

"My phone never rings to interrupt me when I'm spending time with my family."

Trying to ask these as exact numerical questions would often create false precision.


Section 9: Why we retain a neutral midpoint

When a neutral position is meaningful, we include it. Respondents can choose neutral when neither agreement nor disagreement fits.

Survey-design literature generally supports choosing the response structure according to the construct being measured. Where a meaningful neutral position exists, an odd-numbered response scale can be appropriate; where no meaningful midpoint exists, forcing one may be unnecessary. Research also suggests that the consequences of odd versus even response counts are often smaller than designers assume.

In the diagnostic, Neutral is not automatically interpreted as "average performance."

Its score depends on what the question means.

A neutral answer to a belief question may indicate mild exposure.

A neutral answer to a question about confidence may indicate that the business cannot confidently claim the control exists.

We interpret the wording and score together.


Section 10: Frequency questions

Where the construct is recurring behaviour, we often ask directly about frequency.

For example:

How often do jobs become harder, slower or less profitable because the customer was not a good fit from the start?

Possible responses might progress from:

Rarely → Occasionally → Frequently → Almost constantly

This is preferable to asking:

"How much do you agree that bad-fit customers sometimes cause problems?"

This wording asks about the behaviour directly.

Research in survey methodology also shows that response alternatives themselves influence how people interpret behavioural-frequency questions, which is one reason the response ranges need to be chosen deliberately rather than added casually.


Section 11: Percentage bands

Some questions use ranges.

For example:

What percentage of your leads are already pre-qualified when you speak with them for the first time?

Possible answers include:

75% or more 50% to 74% 25% to 49% Less than 25%

The bands capture useful differences without asking respondents for precision they may not have.

A remodeler may reasonably know:

"roughly half"

without maintaining enough historical data to state:

"52.7%."

Forcing the latter would create the appearance of accuracy without necessarily improving the underlying information.


Section 12: Behaviourally anchored response options

Some questions ask respondents to describe what happens instead of judging whether they are "good" or "bad" at something.

They ask what actually happens.

For example:

When another sales conversation is required, what normally happens?

A set of responses can distinguish between:

A firm date and time is booked and centrally recorded

versus:

A casual intention is made to reconnect

versus:

The homeowner is left to contact the company when ready

These options describe recognisably different operating states.

Respondents describe the process, and the diagnostic assigns its severity. Throughout the assessment, we prefer questions about behaviour when we can ask what happens instead of asking respondents to grade themselves.


Section 13: Binary controls

Not every question needs five answer choices.

Some operating controls have a binary answer. For example:

Can you see the total potential value of your current pipeline?

or:

Do you record the outcome of every sales call?

The response is simply:

Yes / No

Intermediate choices can make the question look more detailed while making the answer less clear.

We score those answers directly from low exposure to high exposure.


Section 14: Outcome questions

Where possible, the diagnostic measures both processes and outcomes.

For example, a business may report that it has:

  • strong qualification;
  • documented follow-up;
  • and a controlled sales process.

But if it simultaneously reports:

  • poor conversion;
  • frequent ghosting;
  • or large numbers of unsuitable enquiries,

the difference deserves attention.

Outcome questions let us compare reported results with confidence in the process.

Outcomes still do not prove a particular cause.

A weak conversion rate could be influenced by:

  • lead quality;
  • qualification;
  • price;
  • customer fit;
  • sales execution;
  • competition;
  • or market conditions.

The semantic and cross-category rules help prevent us from assigning a cause without enough evidence.


Section 15: Subjective questions

Some subjective information matters in its own right.

Examples include:

  • confidence that leads are properly worked;
  • stress about future work;
  • confidence in labour availability;
  • confidence when a team member says something is handled.

These questions are useful because management confidence and commercial pressure affect how a business operates.

We do not treat confidence as objective proof.

For example:

A respondent may say:

"I am very confident that every lead is properly worked."

If the same respondent later reports:

"We have no reliable way of recording what happened to each opportunity,"

the diagnostic should not simply accept both statements independently.

The reconciliation rules address this kind of conflict.


Section 16: Reverse-direction questions

The direction of the response scale is not always the same.

Consider these two statements:

"Every new lead gets an instant response from me or my team."

and:

"Most homeowners only care about the lowest price."

Agreement with the first generally indicates stronger control.

Agreement with the second indicates greater exposure.

We score each answer by what it means, not where it appears.

Reverse-worded items also make it harder to complete the assessment by repeatedly choosing answers from one side of the scale.

We do not claim that this makes the diagnostic a formally validated psychometric instrument or that we have independently quantified acquiescence-response bias.

It is a basic questionnaire-design choice.


Section 17: Why five visible answers do not always mean five different scores

A question may show five response choices while two adjacent answers share a severity score.

For example:

ResponseSeverity
Strongly agree0
Somewhat agree1
Neutral2
Somewhat disagree3
Strongly disagree3

Both negative answers receive 3 when they indicate the same operating condition. The scoring does not create a mathematical difference just because the form has five choices.

If both:

Somewhat disagree

and:

Strongly disagree

indicate that an important operational control is substantially absent, inventing a score difference between them would create false precision.

The reverse can also happen.

If both:

Strongly agree

and:

Somewhat agree

represent an acceptably controlled condition, both may map to zero.

The score follows each answer's meaning, not its position in the list.


Section 18: The underlying 0–3 severity model

Every scored answer ultimately maps onto an ordinal severity value.

Let:

sij∈{0,1,2,3}s_{ij} \in \{0,1,2,3\}

where:

  • (i) represents the diagnostic category;
  • (j) represents the question;
  • and (s) represents the severity associated with the selected answer.

The general interpretation is:

ScoreGeneral Meaning
0Little or no meaningful measured exposure
1Limited or occasional exposure
2Material or recurring exposure
3Substantial exposure or absence of an important control

Important: this is an ordinal severity model

A score of 3 does not mean:

"This answer creates exactly three times the financial damage of a score of 1."

A score of 2 does not mean:

"This condition is exactly twice as bad as a score of 1."

The scores create an ordered severity structure.

They are not empirically estimated financial multipliers.

Treating the values as financial multipliers would misrepresent the model.


Section 19: The seven category structures

The final instrument contains 46 scored questions distributed across seven categories.

A purpose record was created for every scored question, and answer-level semantic records were created for every selectable scored response.

CategoryScored QuestionsMaximum Raw Exposure
Lead Capture & Response412
Qualification & Opportunity Prioritisation721
Follow-Up & Sales Progression618
Visibility & Sales Accountability618
Project Selectivity & Customer Fit927
Operational Strain & Owner Dependence927
Delivery & Margin Reality Check515
Total46Not applicable

Categories contain different numbers of questions. Each has enough items to examine its domain; we did not add questions just to make the category sizes match.


Section 20: How a category score is calculated

Layer 1: raw category exposure

For category (c):

Rc=∑j=1ncscjR_c = \sum_{j=1}^{n_c} s_{cj}

where:

  • (R_c) = raw category exposure;
  • (n_c) = number of scored questions in the category;
  • (s_{cj}) = selected answer severity.

Layer 2: maximum possible exposure

Mc=∑j=1ncmcjM_c = \sum_{j=1}^{n_c} m_{cj}

where:

  • (M_c) = maximum possible exposure in the category;
  • (m_{cj}) = maximum possible severity for each question.

In the current diagnostic, every scored question has a maximum severity value of 3.

Therefore, for a four-question category:

Mc=4×3=12M_c = 4 \times 3 = 12

For a nine-question category:

Mc=9×3=27M_c = 9 \times 3 = 27

Layer 3: normalized exposure

Because categories contain different numbers of questions, raw scores cannot be compared directly.

The raw category result is therefore normalized:

Ec=round⁡(RcMc×100)E_c = \operatorname{round} \left( \frac{R_c}{M_c} \times100 \right)

where:

E_c = Normalized Exposure

This produces a value between 0 and 100.


Section 21: Worked example

Consider Lead Capture & Response.

It contains four questions.

Maximum exposure:

4×3=124 \times 3 = 12

Suppose a respondent receives:

1+2+2+1=61+2+2+1=6

Then:

Ec=612×100=50E_c = \frac{6}{12}\times100 = 50

The normalized category exposure is:

50/100 Exposure


Section 22: Why we normalize the categories

Normalization solves one specific mathematical problem.

A category containing four questions cannot be displayed meaningfully beside a category containing nine questions using their raw totals.

For example:

9/12

and:

9/27

are both raw scores of 9.

But they represent completely different positions within their respective categories.

Normalization converts them to:

75/100

and:

33/100

respectively.

Every category is shown on the same 0–100 range.

What normalization does not mean

It does not establish that:

70 points of exposure in Lead Capture has exactly the same financial impact as 70 points in Delivery & Margin.

The common scale makes the category positions easier to understand.

It does not create universal economic equivalence between categories.


Section 23: Why the report gauge runs in the opposite direction

The internal calculation measures exposure.

Higher exposure is worse.

But a graphical gauge is easier to understand if the stronger end appears at the high side of the scale.

The report therefore uses:

Pc=100−EcP_c = 100-E_c

where:

  • (E_c) = exposure;
  • (P_c) = graphical position.

If exposure is:

7575

then graphical position is:

100−75=25100-75=25

The report can therefore display:

Position: 25/100 Exposure: 75/100

No additional analytical meaning is created by this transformation.

It is a display convention.


Section 24: The four interpretation bands

Normalized exposure is grouped into four broad interpretation zones.

ExposureInterpretation
0–24Strong foundation
25–49Some leakage
50–74Significant weakness
75–100Urgent constraint

The bands summarize normalized exposure for readers.

Important limitation

The boundaries are model thresholds.

They were not discovered through a longitudinal study proving, for example, that:

a remodeler scoring 49 behaves materially differently from one scoring 50.

Crossing a boundary changes the report's band; it does not show that business performance changes sharply at that point.


Section 25: Why the diagnostic has no overall score

A simple average of the seven category values would produce one number:

E1+E2+E3+E4+E5+E6+E77\frac{ E_1+E_2+E_3+E_4+E_5+E_6+E_7 }{7}

We could call it Your Remodeler Business Score.

We do not calculate that average.

The system reports seven category results and no overall diagnostic score.

Averaging unrelated operating conditions can hide a serious problem in one area.

Imagine a company with:

  • exceptional Lead Capture;
  • exceptional Qualification;
  • catastrophic Delivery & Margin control.

An overall average could make the company appear reasonably healthy.

Several strong areas do not cancel a delivery problem.

Poor front-end systems should not cancel strong project controls.

The categories describe different parts of the business.

We report them separately.

Formally:

Overall Diagnostic Score is intentionally undefined


Section 26: Why we do not apply category weights

The model does not assign universal weights to the seven categories.

We could have written:

Lead Capture = 25%

Qualification = 20%

Follow-Up = 15%

and so on.

No evidence establishes that those precise weights apply to every remodeling company.

A delivery problem can destroy a highly profitable business.

A lead-response problem can be extremely important for a company buying large volumes of paid leads.

Owner dependence may matter enormously to a business preparing for succession.

Project selection may matter particularly to a company struggling with low-margin filler work.

Their relative economic importance depends on context.

We do not assign universal weights without evidence to support them.


Section 27: Why we do not automatically rank the categories

Severity is not the same thing as strategic priority.

Suppose:

Lead Capture = 80 exposure

and:

Delivery = 60 exposure

It does not logically follow that Lead Capture should be repaired before Delivery.

A true strategic-priority model would also need information such as:

  • economic value affected;
  • number of opportunities affected;
  • implementation cost;
  • difficulty;
  • management capacity;
  • causal dependencies;
  • urgency;
  • expected return;
  • and probability of successful remediation.

The diagnostic does not claim to measure all of those variables.

We considered stronger ranking and prioritisation, then removed it because category severity alone did not justify a universal repair order.

This sets the model's rule:

When the data does not support a stronger conclusion, the model should become less confident rather than more persuasive.


Section 28: Why the numerical score is not the entire diagnosis

A score tells us how much coded exposure exists in a category.

It does not tell us which problems created that exposure.

Consider two hypothetical businesses.

Business A

Lead Capture ItemScore
Response speed0
After-hours coverage0
Conversation rate3
Lead-work visibility3
Total6/12

Normalized exposure:

50/100

Business B

Lead Capture ItemScore
Response speed3
After-hours coverage3
Conversation rate0
Lead-work visibility0
Total6/12

Normalized exposure:

50/100

Both companies receive the same numerical exposure.

They do not have the same problem.

The report interprets individual answers as well as category totals.


Section 29: The answer-level semantic layer

The reporting system contains an interpretation record for every selectable scored answer.

In the completed semantic engine, the 46 scored questions produce 185 answer-level records. The system was designed so every selectable answer could be interpreted individually rather than only contributing points to a total.

For each answer, the system can define:

  • what the respondent actually reported;
  • the severity associated with it;
  • what the answer may reasonably suggest;
  • what it does not establish by itself;
  • which other answers provide relevant context;
  • what finding can legitimately appear;
  • possible operating consequences;
  • and whether a ClientLab capability is potentially relevant.

This is why the report can say:

"You reported X, which means Y deserves investigation…"

rather than simply:

"Your score is orange."


Section 30: Every question had to have a reason to exist

We separated question development from report writing.

Before refining the final report language, each question was reviewed for:

  • its intended purpose;
  • the business issue it was designed to surface;
  • which conclusions it was allowed to support;
  • which conclusions would go beyond the evidence;
  • and whether any ClientLab relationship was genuine.

We then reviewed each answer interpretation by asking:

If a remodeler selected this answer, is this a fair and useful interpretation of what the answer suggests?

This check kept report wording from changing what each question measures.


Section 31: The deterministic interpretation pipeline

The diagnostic follows a defined sequence:

Answer → Severity → Semantic Meaning → Context Rules → Contradiction Checks → Approved Findings → Report

The report follows this sequence without requiring a generative model to interpret answers.

Reproducibility

The same answer profile can be tested repeatedly.

Traceability

A finding can be traced back to the condition that caused it to appear.

Change control

An interpretation cannot silently change because a language model happened to respond differently.

Testability

We can create synthetic companies to test difficult answer combinations.

Reduced hallucination risk

The runtime system cannot simply invent a new explanation that was never reviewed.


Section 32: Contradiction and mixed-signal reconciliation

Self-reported answers can conflict. A person can give answers that make sense individually but are difficult to reconcile together.

For example:

"We respond extremely quickly."

and:

"Evenings and weekends normally wait until somebody becomes available."

Both statements can be true.

The correct conclusion is not necessarily:

"One answer is wrong."

The correct conclusion may be:

"Average working-hours performance appears stronger than coverage outside those periods."

Other contradictions can be more direct.

Examples tested during development included:

  • fast average response but weak evening and weekend coverage;
  • slow actual response combined with a claim of universal instant response;
  • high claimed pre-qualification despite very little information being collected before the first call;
  • budget supposedly being known before the call while the team also reports entering the call completely blind;
  • strict qualification processes that continue producing poor-fit opportunities;
  • documented follow-up systems alongside weak real-world execution;
  • automated reminders that do not appear to protect appointments;
  • strong project-selection claims alongside poor completed-job outcomes;
  • weak selection controls alongside strong recent outcomes;
  • low reported stress alongside substantial schedule or cash-flow pressure;
  • an owner who says they can step away while also reporting excessive oversight;
  • confidence in trade availability despite frequent labour-driven delays;
  • inability to step away despite strong handoff and information answers;
  • complete claimed cost visibility alongside unexplained margin loss;
  • and strong claimed gross-profit performance that cannot be verified through actual project-cost information.

The adversarial test suite deliberately produced 34 conflicting answer combinations across 10 adversarial profiles, and each flagged contradiction was matched to a defined reconciliation rule.


Section 33: Contribution rather than exclusive causation

The report can connect answers, but those connections do not prove exclusive causation.

A typical reasoning chain might be:

Slow or inconsistent response

↓\downarrow

Fewer real conversations

↓\downarrow

Smaller pool of suitable opportunities

↓\downarrow

Less commercial choice

↓\downarrow

Greater pressure to accept weaker-fit work

This is a plausible operating mechanism.

It is not proof that every difficult project inside that company was caused by slow lead response.

Theory-of-change and contribution-analysis methods informed this part of the diagnostic.

We ask:

Does the evidence make this contribution story plausible?

not:

Can we prove that one factor caused everything downstream?

These approaches examine intermediate links and alternative pathways instead of treating a connection between two outcomes as proof of causation.


Section 34: Why the four Business Snapshot questions are unscored

The respondent also provides four pieces of business context:

Annual revenue

Approximate net profit margin

Monthly paid online advertising spend

Average monthly lead or enquiry volume

These questions do not contribute to the seven category scores.

This prevents circular reasoning.

For example:

A lower profit margin should not automatically create a weak Delivery & Margin score.

A high advertising budget should not automatically create a lead-management problem.

Low lead volume should not automatically mean the company has poor sales systems.

The operating questions diagnose the operating conditions.

The Business Snapshot provides context.

We keep the four Business Snapshot questions separate from the 46 scored questions.


Section 35: External research and its role

We use published research to select constructs, explain mechanisms and provide industry or commercial context. It does not set the model's numerical values.

Research can support construct selection

Example:

Lead-response research supports including response speed as an important operating construct.

Research can support mechanism interpretation

Example:

Reminder research supports the proposition that reminders can protect appointment attendance.

Research can provide industry context

Example:

Construction-labour research helps establish why reliable labour availability is a material operating concern.

Research can provide commercial context

Example:

NAHB profitability data helps explain why repeated downstream leakage deserves attention.

What research generally does not do

The external studies did not statistically derive:

  • the exact 0–3 answer scores;
  • the exact 25/50/75 band boundaries;
  • a universal category weight;
  • an overall remodeler-business score;
  • or a dollar value for each point of exposure.

The roles are separate:

Research → Supports why the construct matters

followed by:

Model design → Defines how the construct is represented

followed by:

Testing → Checks whether the resulting interpretations remain coherent


Section 36: Research behind Lead Capture & Response

One of the major evidence sources is the body of Lead Response Management research associated with Dr James Oldroyd and InsideSales/XANT.

The original research examined more than 15,000 web-generated leads and more than 100,000 call attempts.

InsideSales reports that the odds of making successful contact were approximately 100 times greater when attempting contact within five minutes rather than waiting thirty minutes, while the odds of the lead entering the qualification/sales process were approximately 21 times greater.

How we use this evidence

It supports the importance and direction of response speed.

How we do not use it

We do not claim that every remodeling lead behaves identically to the populations in the original study.

We also do not claim that the research mathematically proves our exact response-time scoring boundaries.

The research tells us that response delay is commercially important.

The 0–3 severity mapping remains a modelling choice.


Section 37: Research behind qualification & protecting human sales capacity

Salesforce's State of Sales, Seventh Edition reports that sales professionals spend approximately 40% of the average working week selling and 60% on activities Salesforce classifies as non-selling, including quoting, planning, manual data entry and training.

Accenture has separately described AI-enabled sales systems that enrich and qualify leads, surface relevant information and support sellers before conversations begin. Based on Accenture's reported experience, automated sales support has produced substantial improvements in lead-qualification productivity in some contexts.

How we use this evidence

It supports the principle that skilled human selling capacity is limited and should be protected from work that systems can perform earlier or more consistently.

How we do not use it

We do not claim that a remodeler will automatically recover Salesforce's 60% non-selling allocation or reproduce Accenture's reported improvement figures.

Those populations and implementations are different.


Section 38: Research behind Follow-Up & Sales Progression

Harvard Business Review has published research and analysis connecting formal sales-process discipline with stronger commercial performance.

Reminder research provides a second, different evidence stream.

A Cochrane review of eight randomized controlled trials involving 6,615 participants found that mobile messaging reminders improved appointment attendance compared with no reminder.

Why use healthcare reminder research?

Because the underlying behavioural mechanism is relevant:

  • future commitments are forgotten;
  • circumstances change;
  • attention shifts;
  • reminders re-surface commitments.

Important limitation

A healthcare appointment is not a remodeling sales appointment.

We therefore use this evidence to support the mechanism that reminders can improve attendance.

We do not claim the exact healthcare effect size will be reproduced in remodeling.


Section 39: Research behind Visibility & Sales Accountability

Boston Consulting Group reported that many B2B companies underuse sales data and analytics and may miss 5% to 10% of annual net-revenue uplift as a result.

BCG also links advanced analytics with better lead generation, nurturing and prioritisation.

How we use this evidence

It supports the importance of commercial visibility around:

  • lead source;
  • conversion;
  • progression;
  • prioritisation;
  • salesperson performance;
  • and outcomes.

How we do not use it

A poor ClientLab Visibility score is not converted into a claim that the respondent is losing 5% to 10% of revenue.

The BCG figure establishes that sales intelligence can be economically significant.

It does not calculate the respondent's individual loss.


Section 40: Research behind project selectivity & customer experience

PwC's 2025 Customer Experience Survey reported that:

  • 52% of consumers surveyed had stopped using or buying from a brand because of a bad product or service experience;
  • 29% had stopped because of poor customer experience either online or in person.

How we use this evidence

It supports the principle that customer decisions are not determined by price alone.

The experience created before purchase can influence trust and behaviour.

How we do not use it

The study does not prove how a specific homeowner chooses between two remodeling companies.

It provides broader consumer evidence that experience matters.


Section 41: Research behind operational strain & labour availability

Associated Builders and Contractors estimated that the U.S. construction industry needed to attract approximately 349,000 net new workers in 2026 to meet projected demand.

Its model incorporates inflation-adjusted construction spending, employment, vacancies, unemployment and expected retirements.

How we use this evidence

It supports the importance of labour availability and operational resilience.

Limitation

ABC's data covers the broader construction industry.

It should not be interpreted as a remodeler-specific shortage estimate for every local market.


Section 42: Research behind handoffs, information and rework

Autodesk and FMI's Construction Disconnected research reported that poor project data and miscommunication accounted for 48% of rework on the U.S. construction jobsites studied.

The research attributed approximately:

  • 26% of rework to poor communication;
  • and 22% to poor project information.

How we use this evidence

It supports the relevance of:

  • reliable handoffs;
  • accessible project information;
  • shared operating context;
  • and reducing information trapped in individuals.

Limitation

The research is broader construction evidence and is not a current national remodeling benchmark.


Section 43: Research behind owner dependence and transferability

Gallup reported in 2025 that 74% of employer-business owners in its research planned eventually to sell or transfer ownership of their businesses.

How we use this evidence

It helps establish why owner independence is not merely a lifestyle question.

For many owners, the eventual ability to:

  • install management;
  • transfer the company;
  • pass it to family;
  • or sell it

is a legitimate long-term business objective.

How we do not use it

The diagnostic does not estimate business valuation.

It does not claim that a specific Owner Dependence score causes a specific valuation discount.


Section 44: Research behind delivery & margin control

NAHB's 2026 reporting on its Remodelers' Cost of Doing Business Study found that average remodeler results for 2024 included:

  • 29.9% gross profit margin
  • 6.3% net profit margin

NAHB described the 6.3% average net margin as the highest since 1996.

The large difference between gross and net profit illustrates why apparently modest operational leakages can matter.

How we use this evidence

It provides remodeling-specific context for examining:

  • margin variance;
  • delays;
  • change control;
  • rework;
  • callbacks;
  • and actual-versus-estimated job costs.

How we do not use it

The NAHB average is not a pass/fail threshold.

A remodeler is not scored negatively merely because their own margin differs from the national average.


Section 45: Evidence hierarchy

Sources vary in how directly they apply to residential remodeling, so we group them by relevance.

Level 1: remodeling-specific evidence

Examples:

  • NAHB remodeler profitability data.

This has the strongest direct population relevance.


Level 2: wider construction evidence

Examples:

  • ABC workforce research;
  • Autodesk/FMI rework and communication research.

This is directly relevant to construction operations but does not exclusively represent residential remodeling.


Level 3: adjacent sales and commercial evidence

Examples:

  • Salesforce;
  • Boston Consulting Group;
  • Accenture;
  • Harvard Business Review;
  • Lead Response Management research.

These support general commercial mechanisms.


Level 4: cross-domain mechanism evidence

Example:

  • healthcare appointment-reminder research.

This can support a behavioural mechanism while requiring substantially more caution about direct effect sizes.


Section 46: Research-to-model map

SourceMain Finding or Principle UsedWhere It Informs the DiagnosticWhat It Does Not Establish
Lead Response Management researchResponse delay substantially reduces contact/qualification oddsLead Capture & ResponseExact remodeler conversion loss or exact 0–3 score boundaries
Salesforce State of SalesLarge share of salesperson time is consumed outside direct sellingQualification, Sales Capacity, VisibilityExact hours every remodeler will recover
Accenture Digital Inside SalesAI/automation can support lead enrichment and qualificationQualification architectureGuaranteed productivity gain for ClientLab customers
HBR formal sales-process researchFormal sales process is associated with stronger sales management/performanceFollow-Up & Sales ProgressionExact revenue gain caused by documentation
Cochrane reminder reviewReminders can improve appointment attendanceConfirmations and appointment protectionExact remodeling no-show reduction
BCG sales-intelligence researchBetter use of sales data and analytics can materially affect performanceVisibility & AccountabilityIndividual respondent's revenue loss
PwC Customer Experience SurveyCustomer experience affects customer retention and buying behaviourCustomer Fit / ExperienceExact homeowner contractor-selection behaviour
ABC workforce modelConstruction faces significant labour demandOperational StrainIdentical labour conditions in every remodeler's market
Autodesk/FMIPoor information and communication contribute materially to reworkHandoffs, Owner Dependence, DeliveryCurrent remodeler-specific rework percentage
Gallup succession researchMany employer owners ultimately plan to sell or transferOwner DependenceIndividual business valuation
NAHB Cost of Doing BusinessRemodeling margins leave limited room for repeated leakageDelivery & MarginA universal target margin for every company

Section 47: Published evidence vs model decisions

We distinguish four kinds of evidence and interpretation:

Published evidence

A finding reported by an external source.

Example:

Lead-response research reports a sharp decline in contact odds as response time increases.


Respondent evidence

Information provided by the remodeler.

Example:

"We normally respond within one hour."


Model decision

A rule ClientLab had to define.

Example:

"Within one hour" receives an ordinal severity value within the Lead Capture category.


Derived output

A calculation produced from the model.

Example:

7/12=58% normalized exposure7/12 = 58\%\text{ normalized exposure}

These should never be confused.

We label model assumptions as modelling decisions so readers do not mistake them for published findings.


Section 48: How we reduced the risk of confirmation bias

ClientLab developed the diagnostic and sells systems intended to address some of the problems it examines. That creates a potential conflict of interest. We disclose the conflict and use controls to reduce its effect on the diagnostic.

48.1 The diagnostic was not built from the ClientLab feature list

The operating journey was defined first.

Product mapping came later.


48.2 Not every category points to ClientLab

The Delivery & Margin Reality Check intentionally examines problems such as:

  • job costing;
  • change-order control;
  • project delays;
  • callbacks;
  • and rework.

ClientLab does not claim to solve those problems directly.

Including these problems makes it less likely that every finding will lead to:

"Buy our software."


48.3 Strong results are allowed to remain strong

The category logic preserves strong results instead of manufacturing a problem in every section.


48.4 Product relevance is evaluated after diagnosis

The ClientLab bridge comes after the diagnosis. A product mapping cannot change the diagnostic finding, and changing a ClientLab capability does not change that finding.


48.5 Unsupported ranking was removed

A more aggressive model could have ranked every weakness and declared an automatic repair order.

The available data did not support that certainty, so we removed the ranking logic.


48.6 Contradictory evidence is not ignored

The system does not simply use whichever answer supports the most dramatic conclusion.

Mixed evidence can weaken, qualify or change a finding.


48.7 Business size does not determine operational score

Revenue, advertising spend, lead volume and reported margin are kept outside the category calculations.


48.8 Runtime AI is not allowed to invent a diagnosis

Interpretation is deterministic.


Section 49: Structured review before implementation

We reviewed the diagnostic in separate layers so decisions about one part would not silently change another.

The review sequence separated:

Assessment structure

Question wording, answer wording, scoring and bands.

Question purpose

Why each question existed and what it was allowed to establish.

Answer semantics

Whether every possible selected answer received a fair interpretation.

Category interpretation

Desired state, findings, consequences and improvement direction.

Cross-category logic

Whether multiple answers actually justified a combined conclusion.

ClientLab mapping

Whether a diagnosed problem had a genuine ClientLab connection and what the boundary of that connection was.

Report assembly

Whether the final combination remained coherent.

Test businesses

Whether realistic and deliberately difficult answer profiles produced sensible output.

We documented the review so later changes could be traced through the parts of the system they affect.


Section 50: Validation and adversarial testing

We tested the semantic engine with internally consistent profiles and profiles containing difficult, conflicting answers.

The final validation suite rerendered 27 complete profiles, including:

  • 10 adversarial random profiles;
  • 4 coherent band profiles;
  • 3 realistic coherent profiles;
  • and 10 earlier mixed/regression profiles.

The same test set exercised all 185 answer records and all 46 scored questions.

Structural checks included

  • every scored answer mapped correctly;
  • no unmapped answer records;
  • no unresolved rule keys;
  • no runtime writing or interpretation;
  • correct report assembly;
  • and regression consistency after rule changes.

Section 51: Contradiction testing

The ten adversarial profiles produced 34 deliberately conflicting answer combinations.

Every identified contradiction was required to map to a defined reconciliation rule.

The test asked how the model handles conflicting answers from respondents.


Section 52: Testing was allowed to prove the model wrong

Testing had to be able to reduce the model's confidence.

For example, the final review corrected two overstatements.

First:

project-selection pressure had previously been treated too readily as evidence that poor-fit work was already entering production.

That was changed.

A company can feel pressure while still successfully protecting its standards.

Second:

the product bridge was narrowed so a company demonstrating strong lead response and qualification would not receive a complete SmartCloser remediation narrative merely because ClientLab offered those capabilities.

The test should confirm that the model runs and identify conclusions it states too confidently.


Section 53: Why we call the result "exposure"

We use exposure to describe the operating conditions indicated by a respondent's answers.

A high exposure score does not necessarily mean the company is currently failing.

A profitable business can still contain weak systems.

Strong demand, talented employees, extraordinary owner effort or favourable market conditions can temporarily compensate for operational weaknesses.

Likewise, a green result does not mean nothing can go wrong.

The score describes how strongly the answers indicate certain operating conditions. It is not a direct measure of company performance.


Section 54: What a green result means

A Green result indicates that the category appears predominantly controlled based on the available answers.

It does not mean:

"This area is perfect."

Because category scores aggregate multiple questions, a company can still have an isolated weakness inside an otherwise strong category.

For example, in a nine-question category a single severity-3 answer contributes:

327×100=11.1%\frac{3}{27}\times100 = 11.1\%

That category could remain green.

Answer-level interpretation keeps important individual responses visible alongside the aggregate.


Section 55: What a red result means

A Red result means the respondent has accumulated a high proportion of the category's maximum coded exposure.

It supports a statement such as:

"Multiple answers indicate substantial exposure in this part of the business."

It does not establish:

"This is the sole cause of your profitability problem."

A red result describes coded exposure; it does not identify a sole cause.


Section 56: What the diagnostic can reasonably say

The diagnostic can make statements such as:

Your answers indicate substantial exposure in Lead Capture & Response.

Follow-up appears to depend heavily on manual memory or disconnected reminders.

Management currently lacks enough information to distinguish lead-source problems from sales-process problems.

The reported level of pre-qualification conflicts with how little information appears to be available before the first conversation.

Project-selection pressure is present.

Normal operations appear heavily dependent on the owner.

Actual project economics are not consistently matching what was expected when work was sold.

Each statement is tied to the respondent's answers.


Section 57: What the diagnostic does not claim

The diagnostic does not claim that:

A score of 75 means the company is worse than 75% of remodelers.

There is currently no representative remodeler population norm behind the 0–100 scale.


It does not claim that:

A score of 3 causes exactly three times the financial loss of a score of 1.

The scale is ordinal.


It does not claim that:

A Red category will cause a specific amount of lost profit.

The diagnostic does not calculate causal dollar losses from category scores.


It does not claim that:

The highest-exposure category should always be repaired first.

Severity is not identical to strategic priority.


It does not claim that:

ClientLab fixes every weakness in the report.


It does not claim that:

Research conducted in B2B sales, healthcare or general construction produces identical effect sizes in residential remodeling.

Adjacent research is used carefully for mechanisms and context.


It does not claim that:

The respondent's answers have been independently verified.

They generally have not.


Section 58: Important measurement limitations

58.1 Self-reporting

The assessment depends on information supplied by the respondent.

Possible sources of error include:

  • imperfect memory;
  • optimism;
  • pessimism;
  • social-desirability bias;
  • inconsistent internal data;
  • misunderstanding a question;
  • or intentional misrepresentation.

Contradiction rules can identify some tensions.

They cannot eliminate all self-report error.


58.2 Ordinal values are added together

The 0–3 system creates ordered severity categories.

Adding those values assumes that the sum provides a useful representation of overall exposure within that domain.

That is a practical modelling assumption.

It is not proof that the psychological or financial distance between every adjacent score is identical.


58.3 No population percentile

The 0–100 result is:

percentage of maximum possible coded exposure

It is not:

percentile compared with other remodelers


58.4 Category bands are interpretive thresholds

The four bands aid communication.

They are not empirically proven performance cliffs.


58.5 Equal contribution within a category is still a modelling assumption

Most questions can contribute up to three severity points.

That does not prove every severity-3 answer has identical real-world economic impact.


58.6 No formal psychometric validation is currently claimed

The diagnostic has not been presented as a psychometrically validated clinical or personality instrument.

The current methodology does not claim published:

  • factor-analysis validation;
  • Cronbach's alpha;
  • McDonald's omega;
  • test-retest reliability;
  • nationally representative norms;
  • or predictive-validity coefficients.

Those would require a different type and scale of empirical validation programme.


Section 59: What formal empirical validation could look like in the future

A future validation programme could recruit a sufficiently large sample of remodeling businesses and combine diagnostic responses with independently verified outcomes.

Potential comparison data could include:

  • actual response times;
  • contact rates;
  • qualified appointment rates;
  • appointment attendance;
  • sales conversion;
  • lead-source conversion;
  • sales activity;
  • project gross margins;
  • margin variance;
  • rework;
  • callbacks;
  • schedule variance;
  • owner hours;
  • and business outcomes over time.

That would make it possible to test:

Content validity

Do independent remodeler and operations experts agree that the questions adequately represent the seven domains?

Construct validity

Do observed answer patterns support the proposed category structure?

Convergent validity

Do questionnaire answers correspond with independently observed operating data?

Criterion validity

Do category scores relate meaningfully to relevant business outcomes?

Predictive validity

Do current diagnostic patterns help predict future operational outcomes?

Test-retest reliability

Does a stable company receive a reasonably stable result when nothing meaningful has changed?

Responsiveness

Does the diagnostic detect verified improvement after a business actually changes its operating process?

Such research could eventually justify recalibration of:

  • question thresholds;
  • severity values;
  • category boundaries;
  • or category structure.

Until that evidence exists, we do not present the current numerical model as empirically fitted to population outcomes.


Section 60: Complete calculation framework

For reproducibility, the essential scoring model can be stated compactly.

Answer severity

scj∈{0,1,2,3}s_{cj} \in \{0,1,2,3\}

Raw category exposure

Rc=∑j=1ncscjR_c = \sum_{j=1}^{n_c}s_{cj}

Maximum category exposure

Mc=∑j=1ncmcjM_c = \sum_{j=1}^{n_c}m_{cj}

Normalized exposure

Ec=round⁡(100RcMc)E_c = \operatorname{round} \left( 100\frac{R_c}{M_c} \right)

Graphical position

Pc=100−EcP_c = 100-E_c

Overall score

Not calculated


Category weighting

Not applied


Automatic category ranking

Not applied


Section 61: Worked example: same score, different diagnosis

This example shows why answer-level interpretation matters.

Remodeler A

Response speed: 0 After-hours coverage: 0 Conversation rate: 3 Lead-work visibility: 3

0+0+3+3=60+0+3+3=6 6/12=50%6/12=50\%

Remodeler B

Response speed: 3 After-hours coverage: 3 Conversation rate: 0 Lead-work visibility: 0

3+3+0+0=63+3+0+0=6 6/12=50%6/12=50\%

Both companies receive:

50/100 exposure

But Remodeler A appears to have a conversation/visibility problem despite strong response mechanics.

Remodeler B appears to have a response-coverage problem despite stronger reported outcomes and visibility.

The report preserves this distinction by interpreting each answer alongside the category score.


Section 62: Summary of the methodology

The ClientLab Remodeler Diagnostic is designed around several principles.

  • Measure the operating journey, not the ClientLab product.

  • Use the question format that fits the construct rather than forcing every question into one template.

  • Use ordered severity scores without pretending they are precise financial effect sizes.

  • Normalize categories independently because they contain different numbers of questions.

  • Do not collapse seven different business domains into one misleading overall score.

  • Do not invent category weights simply to make the model look sophisticated.

  • Do not assume the highest-exposure category automatically deserves first strategic priority.

  • Interpret the actual answers, not merely the score total.

  • Reconcile contradictions rather than pretending self-reported data is perfectly consistent.

  • Use contribution reasoning rather than unsupported claims of exclusive causation.

  • Separate diagnosis from product mapping.

  • Make commercial conflicts visible rather than pretending they do not exist.

  • Use published research to support constructs and mechanisms without pretending external studies calibrated numbers they did not calibrate.

  • Test difficult profiles alongside ordinary ones.

  • Allow testing to reduce or remove conclusions when the evidence does not support them.


Section 63: Closing statement

ClientLab sells products related to some problems the diagnostic measures. That creates a risk that the findings could be shaped to support a sale. A diagnostic used as a sales tool can overstate the number or urgency of problems, connect each problem to the product, or present uncertain findings too confidently. We designed the methodology to limit those risks.

The diagnostic includes areas ClientLab does not solve.

Strong results are allowed to remain strong.

Context questions do not manipulate category scores.

Categories are not weighted or combined without evidence.

Contradictory answers can weaken conclusions.

Product mapping occurs after the diagnosis.

External research is separated from our own modelling assumptions.

When testing showed that an interpretation went beyond what the answers supported, we revised it.

The diagnostic is not perfect, but it makes its assumptions visible. We believe every serious business diagnostic should do the same.


Appendix A: Business Snapshot questions

These questions provide context and do not affect category scoring.

A1. Annual revenue

Used to understand the approximate scale of the business.

A2. Approximate net profit margin

Used to understand reported economic context.

A3. Monthly paid online advertising spend

Used to understand current investment in paid demand generation.

A4. Average monthly leads or enquiries

Used to understand approximate lead volume.


Appendix B: Complete scored question framework

Category 1: Lead Capture & Response

Q1. First real response speed

Instantly: 0 Within 15 minutes: 1 Within 1 hour: 2 More than 1 hour: 3


Q2. Universal instant-response coverage

Strongly agree: 0 Somewhat agree: 1 Neutral: 2 Somewhat disagree: 3 Strongly disagree: 3


Q3. Percentage becoming meaningful two-way conversations

Over 75%: 0 51–74%: 1 26–50%: 2 Less than 25%: 3


Q4. Confidence that leads are properly worked

Very confident: 0 Somewhat confident: 1 Not very confident: 2 No real way to track this: 3


Category 2: Qualification & Opportunity Prioritisation

Q5. Information available before the first conversation

Detailed project parameters and confirmed budget bracket: 0 Basic contact information and general room description: 1 Name and phone/email only: 2 Completely blind until the call starts: 3


Q6. Initial qualification process

Strict documented checklist followed by everyone: 0 Loose set of questions generally asked: 1 Informal conversation based on intuition: 2 No standardised process: 3


Q7. When budget is first discussed

Known before the first conversation: 0 Asked during the first short sales conversation: 1 Usually discovered during the home visit: 2 Not normally asked upfront: 3


Q8. Homeowner openness about budget

Strongly agree: 0 Somewhat agree: 1 Neutral: 2 Somewhat disagree: 3 Strongly disagree: 3


Q9. Percentage pre-qualified before the first conversation

75% or more: 0 50–74%: 1 25–49%: 2 Under 25%: 3


Q10. Incoming leads represent the ideal client

Strongly agree: 0 Somewhat agree: 0 Neutral: 1 Somewhat disagree: 2 Strongly disagree: 3


Q11. Time-wasting or non-serious enquiries

Rarely: 0 Occasionally: 1 Most weeks: 2 Daily: 3


Category 3: Follow-Up & Sales Progression

Q12. How follow-up is remembered

Automated reminders: 0 Note in phone: 1 Paper note: 2 Memory: 3


Q13. Reminders for scheduled sales calls

Every time: 0 Usually: 1 Occasionally: 2 Hardly ever / never: 3


Q14. Ghosting after demonstrated interest

Rarely: 0 Occasionally: 1 Frequently: 2 Almost always: 3


Q15. How the next sales conversation is arranged

Firm date/time booked immediately and logged centrally: 0 Casual note to reconnect later: 1 Prospect is left to contact the company when ready: 3


Q16. Documentation of the sales / qualification / follow-up process

Fully documented and followed: 0 Loose notes for parts: 1 Very little documented: 2 Nothing documented / everybody operates independently: 3


Q17. Estimate-to-signed-job conversion

Over 75%: 0 51–75%: 1 26–50%: 2 0–25%: 3


Category 4: Visibility & Sales Accountability

Q18. Knowledge of conversion rate by lead source

Strongly agree: 0 Somewhat agree: 1 Neutral: 2 Somewhat disagree: 3 Strongly disagree: 3


Q19. Knowledge of total potential pipeline value

Yes: 0 No: 3


Q20. Exact cost per lead by acquisition channel

Exact dollar figure: 0 Rough idea requiring manual calculation: 1 Only overall marketing spend is tracked: 3


Q21. Salesperson conversion visibility

Clear daily-updated dashboard: 0 Available through manual calculation: 1 Mostly gut feel: 2 No conversion tracking: 3


Q22. Lost-sale reason tracking

Standardised report for every lost sale: 0 Manual notes or spreadsheets: 1 Ask salesperson / infer from memory: 2 No meaningful tracking: 3


Q23. Recording outcomes of every sales call

Yes: 0 No: 3


Category 5: Project Selectivity & Customer Fit

Q24. Estimates used for shopping around

Rarely: 0 Sometimes: 1 Frequently: 2 Almost every time: 3


Q25. "Most Homeowners Only Care About the Lowest Price"

Strongly disagree: 0 Somewhat disagree: 0 Neutral: 1 Somewhat agree: 2 Strongly agree: 3


Q26. Ability to walk away from a bad-fit project

Strictly protect standards and walk away easily: 0 Usually, with occasional exceptions: 1 Only when warning signs are extreme: 2 Rarely because passing up revenue feels too risky: 3


Q27. Accepting jobs mainly to protect cash flow

Never: 0 Occasionally: 1 Often: 2 Constantly: 3


Q28. Frequency of having to justify pricing

Rarely: 0 Sometimes: 1 Frequently: 2 Almost every time: 3


Q29. Poor customer fit making jobs harder, slower or less profitable

Rarely: 0 Occasionally: 1 Frequently: 2 Almost constantly: 3


Q30. Last ten jobs the business would gladly repeat

9–10: 0 7–8: 1 4–6: 2 0–3: 3


Q31. Pressure to keep feeding the machine

Healthy margins/control permit selectivity: 0 Pressure exists but is manageable: 1 Constant pressure to keep the schedule full: 2 A slowdown feels dangerous and work is needed almost regardless of fit: 3


Q32. Stress about where future jobs will come from

No stress / clear pipeline: 0 Not very stressed: 1 Neutral: 2 Somewhat stressed: 3 Very stressed: 3


Category 6: Operational Strain & Owner Dependence

Q33. Access to preferred trades, crews and subcontractors

Very confident / treated as a priority: 0 Moderately confident: 1 Not very confident: 2 Completely uncertain: 3


Q34. Delays because required people are unavailable

Rarely: 0 Occasionally: 1 Frequently: 2 Constantly: 3


Q35. Ability to step away for a full week

Easily: 0 Usually, with minimal issues: 1 Significant friction or firefighting: 2 Absolutely not: 3


Q36. Being the only person with the full picture

Rarely: 0 Sometimes: 1 Most weeks: 2 Daily: 3


Q37. Confidence When Somebody Says "It's Handled"

Fully confident: 0 Mostly confident; verify major milestones: 1 Skeptical; details are regularly missed: 2 Must micromanage: 3


Q38. "My Phone Never Rings to Interrupt Me When I'm Spending Time With My Family"

Strongly agree: 0 Somewhat agree: 1 Neutral: 2 Somewhat disagree: 3 Strongly disagree: 3


Q39. Personally chasing updates

Rarely: 0 A few times a week: 1 Most days: 2 Constantly: 3


Q40. Weak handoffs creating avoidable problems

Rarely: 0 Occasionally: 1 Often: 2 Constantly: 3


Q41. Effect of growth on time, stress and freedom

Business needs less of the owner / more control: 0 Some improvement but owner remains heavily involved: 1 More work and pressure than previously: 2 Business cannot operate properly without owner: 3


Category 7: Delivery & Margin Reality Check

Q42. Actual gross profit versus expected gross profit

Close on almost every project: 0 Close on most projects: 1 Close on some projects: 2 Rarely, or no reliable way to know: 3


Q43. Estimated-versus-actual project cost visibility

Complete comparison for every project: 0 Usually possible with some manual work: 1 Information is incomplete or inconsistent: 2 Cannot reliably compare: 3


Q44. Work beginning before documented change-order approval

Never: 0 Occasionally: 1 Frequently: 2 Constantly: 3


Q45. Delays reducing profit or blocking the next planned start

Rarely: 0 Occasionally: 1 Frequently: 2 Constantly: 3


Q46. Rework or callbacks consuming unbudgeted resources

Rarely: 0 Occasionally: 1 Frequently: 2 Constantly: 3


Appendix C: Research source register

Lead Response Management Study / InsideSales / XANT

Used to support the importance and direction of response speed.

Lead Response research summary


Salesforce: State of Sales, Seventh Edition

Used principally to support the sales-capacity and administrative-work discussion.

Salesforce State of Sales report


Boston Consulting Group: What If B2B Companies Trusted Their Sales Intelligence?

Used to support the importance of sales visibility, analytics and prioritisation.

BCG sales-intelligence research


Accenture: Elevate Every Decision With Digital Inside Sales

Used as adjacent evidence around automated sales support, lead enrichment and qualification.

Accenture digital inside sales research


Harvard Business Review: Companies With a Formal Sales Process Generate More Revenue

Used to support the relevance of repeatable sales-process discipline.

Harvard Business Review article


Cochrane: Mobile Phone Messaging Reminders for Attendance at Healthcare Appointments

Used as cross-domain evidence for the behavioural mechanism behind appointment reminders.

Cochrane review


PwC: 2025 Customer Experience Survey

Used to support the proposition that customer experience can materially influence consumer behaviour.

PwC Customer Experience Survey


Associated Builders and Contractors: Construction Workforce Shortage Model

Used to provide construction-industry context around labour availability.

ABC 2026 workforce estimate


Autodesk / FMI: Construction Disconnected

Used to support the importance of reliable information, communication and handoffs.

Autodesk/FMI construction research


Gallup: Most Small-Business Owners Lack a Succession Plan

Used to support the relevance of owner independence and business transferability.

Gallup succession research


NAHB: Remodelers' Cost of Doing Business

Used to provide remodeling-specific economic context for Delivery & Margin.

NAHB remodeler profitability summary


Appendix D: Methodological references

Survey and response-scale design

The diagnostic's use of mixed response formats draws on general survey-design principles rather than assuming that every construct should be measured using one identical scale.

Survey-method research shows that:

  • response alternatives influence answers;
  • different constructs warrant different response formats;
  • and meaningful midpoints can be appropriate where a genuine neutral position exists.

Likert-type measurement

A distinction is made between a complete Likert scale and an individual Likert-type item.


Criterion-referenced assessment

The diagnostic compares reported operating conditions with defined criteria rather than population percentiles.


Theory of change and contribution reasoning

Cross-category interpretation was influenced by the broader methodological principle of examining plausible causal pathways and contributory factors without assuming exclusive causation.


Appendix E: Independence and conflict-of-interest statement

The ClientLab Remodeler Diagnostic was developed by ClientLab.

It is therefore not independent third-party research.

ClientLab has a commercial interest in some of the operating problems examined by the diagnostic because ClientLab sells systems intended to improve parts of the remodeling customer-acquisition and pre-sales process.

We believe that should be disclosed rather than obscured.

To reduce the risk that commercial interests determine diagnostic conclusions, the methodology uses:

  • separate diagnostic and product-mapping layers;
  • areas outside ClientLab's direct product scope;
  • answer-level semantic rules;
  • strong-result protection;
  • contradiction reconciliation;
  • deterministic report generation;
  • unscored commercial-context questions;
  • removal of unsupported category ranking;
  • adversarial test profiles;
  • regression testing;
  • and documented review of conclusions that testing showed were too strong.

These controls reduce bias risk.

They do not make ClientLab an independent third party.


Appendix F: Final interpretation rule

The simplest way to understand the entire methodology is:

Ask what actually happens

↓\downarrow

Code the severity without pretending it is a precise economic effect

↓\downarrow

Interpret the actual answer, not merely the score

↓\downarrow

Check the surrounding evidence

↓\downarrow

Reconcile contradictions

↓\downarrow

Identify plausible operating consequences without overstating causation

↓\downarrow

Only then determine whether ClientLab has anything relevant to offer

That is the methodology behind the ClientLab Remodeler Diagnostic.