Skip to content
Searcle Book a demo

Build an Explainable Buyer-Qualification Workflow in Clay

Nina Okonkwo

Buyer scoring is useful when it supports a consistent go-to-market decision: who sales should work now, who should enter nurture, which accounts deserve named ownership, and which records need more research.

It becomes less useful when every attribute is compressed into an unexplained number. A score of 78 tells an operator little unless they can see whether it came from strong customer fit, recent engagement, repetitive low-value activity, penalties, or incomplete data.

A practical Clay workflow should therefore preserve its components. Keep fit, engagement, and negative signals visible; record the source and freshness of important data; calculate scores through transparent Formula columns; and connect qualified tiers to explicit CRM actions. Treat every weight and threshold as a hypothesis to be tested against actual sales outcomes.

1. Define what the Clay buyer score is supposed to decide

A buyer score is a prioritization mechanism. It ranks or classifies people and accounts under criteria your organization defines. It should not be interpreted as a precise probability that someone will purchase.

A high score can justify faster research, sales assignment, or a nurture change without claiming that the record will become a customer. The score expresses how closely the available evidence matches your current qualification rules.

Before choosing attributes, write down the decision the score will control. Common decisions include:

  • Prioritizing inbound leads
  • Assigning records to sales representatives
  • Routing strategic accounts to named owners
  • Placing lower-priority records in nurture
  • Suppressing competitors, students, or irrelevant contacts
  • Sending incomplete or contradictory records to manual review
  • Selecting accounts for outbound research
  • Changing lifecycle stages or campaign membership

A model built for inbound response will not necessarily work for account-based outbound. Inbound scoring may emphasize recent requests and form activity. Account selection may rely more heavily on firmographics, technologies, hiring, and buying-group coverage. Product-led qualification may require a separate product-activity component.

Keep fit and engagement distinct

Clay distinguishes customer or account fit, lead or account engagement, and a combined lead grade. Its examples include company size, industry, and technology stack for fit, and email responses, website visits, and event participation for engagement. Users define the data points and criteria relevant to their businesses in Clay’s lead-scoring documentation.

That does not mean the components should immediately be collapsed.

Fit asks whether the person or account resembles the customers the organization is prepared to serve. It may include industry, company size, geography, technology, role, and seniority. These attributes usually change relatively slowly.

Engagement describes observed activity, such as a demo request, pricing-page visit, email response, event participation, or product use. Selected high-value engagement and timing signals may serve as intent proxies, but activity does not by itself prove purchase intent. These signals can change quickly and may lose relevance over time.

Keeping the components visible prevents several errors:

  • A poor-fit account cannot qualify solely through passive activity.
  • A strong-fit account is not mistaken for an active buyer.
  • Sales can distinguish “right company, quiet now” from “active contact, questionable fit.”
  • Engagement can be refreshed without unnecessarily recomputing every fit attribute.
  • Analysts can identify whether routing failures came from the ICP definition or behavioral rules.

A final priority score may combine the components, but the original scores should remain available for inspection.

Use an explainable scoring schema

A practical implementation framework contains at least these fields:

Field Purpose
Fit Score Measures person and account alignment with the ICP
Engagement Score Measures eligible recent activity and timing signals
Negative Score Stores the nonnegative magnitude of applicable penalties
Final Score Combines fit and engagement, then subtracts penalties
Tier Converts the result into an operational category
Score Reason Summarizes the decisive positive and negative rules
Data Confidence Indicates whether the underlying data is reliable enough
Data Freshness Shows when important attributes were last verified
Next Action States what should happen to the record

This is an editorial implementation framework, not an official Clay template. Adapt it to your sales motion, CRM structure, and data environment.

For every qualified tier, define:

  1. The destination or owner
  2. The action to take
  3. The reason shown to the recipient
  4. The fields written to the CRM
  5. The conditions that suppress or override the action
  6. The fallback path for missing or contradictory data

The formula becomes useful only when sales, marketing, and operations agree on what its output means.

2. Choose person-level, account-level, intent, and negative attributes

Do not begin with one undifferentiated list of fields. Organize the attribute dictionary by the level of the signal and the purpose it serves.

A practical taxonomy includes person-level fit, account-level fit, engagement, timing, and negative or disqualification signals.

Person-level fit attributes

Person-level attributes evaluate whether the contact is relevant to the buying process:

  • Job title: Raw title and normalized title category
  • Function or department: Sales, finance, operations, IT, marketing, security, or another relevant function
  • Seniority: Individual contributor, manager, director, vice president, executive, or owner
  • Buying role: Decision-maker, champion, evaluator, user, influencer, procurement participant, or unknown
  • Geography: Contact location when it affects territory, language, availability, or eligibility
  • Contact validity: Whether email, phone, and employment details appear usable and current

Title alone is rarely sufficient. A director in an irrelevant department may deserve fewer points than a manager who owns the problem your product addresses. A target title should not receive full credit when the person has left the company or the contact information is invalid.

Buying-role classifications can improve prioritization, but inferred classifications should be treated as uncertain. Preserve their source and confidence.

Account-level fit attributes

Account-level attributes describe whether the organization fits the target market:

  • Industry or subindustry
  • Employee or revenue band
  • Country, region, or sales territory
  • Technology stack
  • Funding status
  • Company growth indicators
  • Relevant open roles or hiring patterns
  • Ownership or business model, where relevant
  • Existing customer, partner, competitor, or target-account status

Use ranges where exact values are unreliable or unnecessary. An employee band may be enough for routing even if sources disagree about exact headcount.

Firmographic fit should reflect practical constraints. A company may resemble successful customers statistically while being outside supported regions, too small for the implementation model, or already owned by another sales team. Score the market the organization can actually serve.

Technographic attributes

Technographic signals may include:

  • Use of a complementary technology
  • Use of a competing product
  • Absence of a prerequisite technology
  • A recent technology adoption or removal
  • Use of a category associated with the target problem

A competing technology can be positive, negative, or neutral. It may indicate a replacement opportunity, contractual lock-in, or an incompatible architecture. Its treatment depends on the sales motion.

Technology changes can also raise research priority, but they do not prove active demand. Retain a last-verified date.

Behavioral and engagement attributes

Possible engagement attributes include:

  • Demo or consultation requests
  • Pricing-page visits
  • Repeat website sessions
  • Form submissions
  • Case-study or guide downloads
  • Webinar registration and attendance
  • Email replies or meaningful clicks
  • Product signups, logins, feature use, or workspace activity
  • Chat conversations
  • Engagement from several relevant contacts at one account

These actions are not equally informative. A direct demo request is usually a clearer request for interaction than an email open. A pricing-page visit may be more commercially relevant than a general blog view. Engagement across a buying group may be more useful than repeated low-value activity from one irrelevant contact.

The eventual weights should come from the organization’s own conversion and sales-acceptance data. An action that appears important in theory may not correlate with useful opportunities in a particular funnel.

Passive signals need special caution. Such activity can contribute modestly, but it should be capped and should not override poor fit.

Timing signals

Timing signals can identify accounts worth investigating:

  • A recent funding event
  • Hiring for roles connected to the problem
  • Rapid employee growth
  • Leadership changes
  • Adoption or removal of relevant technology
  • Expansion into a new region
  • A legitimately available renewal window
  • A recent product or strategic announcement

These signals provide context rather than proof of demand. Funding may create budget or precede cost controls. Hiring may indicate investment in a function or a plan to solve the problem internally.

Use timing signals to strengthen other evidence or raise research priority, not as definitive purchase intent.

Negative and disqualification attributes

Negative rules help prevent misleading positive signals from dominating the model. Candidate attributes include:

  • Competitor status
  • Student or academic-research status
  • Irrelevant role or department
  • Excluded industry or region
  • Invalid, bounced, or personal email address
  • Unsubscribe status
  • Career-page activity
  • Existing customer or open opportunity
  • Duplicate record
  • Prolonged inactivity
  • Explicit sales disqualification
  • Evidence that the person no longer works at the account

Negative scoring can create false negatives. A personal email may belong to a legitimate founder. A career-page visit may come from a buyer researching the company. Inactivity may reflect a long buying cycle.

Missing data is not disqualifying evidence. “Industry unknown” does not mean “wrong industry,” and “technology not detected” does not mean the technology is absent. Award no points, mark the record incomplete, or request enrichment rather than automatically applying a penalty.

Build an attribute dictionary

Document every candidate signal before writing the formula:

Dictionary field Question answered
Signal name What is being evaluated?
Level Does it belong to a person or account?
Purpose Is it fit, engagement, timing, or negative evidence?
Source system Where does the value originate?
Normalized values Which controlled categories will the formula use?
Proposed weight How many points or what condition applies?
Cap Can repeated events inflate the score?
Decay rule Does the signal lose value over time?
Refresh cadence How often should it be rechecked?
Confidence requirement What quality is required before awarding points?
Business owner Who approves the definition and future changes?

This dictionary becomes the contract between the formula and the business process. It also exposes rules that cannot be implemented because the required data is unavailable.

3. Map every signal to a source and add data-quality controls

Create a source map before writing formulas. Every scoring rule should correspond to an available field, event, or enrichment output.

For each attribute, identify:

  • The originating system
  • The Clay column that receives it
  • The identifier used for matching
  • Whether the value is raw, normalized, enriched, inferred, or aggregated
  • How often it can change
  • How conflicts are resolved
  • What happens when it is unavailable

Separate ingestion from signal collection

Records can enter a Clay workflow through patterns such as CSV uploads, CRM synchronization, webhooks, forms, Google Sheets, searches, and other imports. These patterns are described in a third-party Clay implementation guide, but connector availability and configuration should be checked in the current product interface.

Profile and company information may then be imported or enriched, including title, seniority, industry, employee band, location, revenue band, and technology.

Behavioral events usually originate elsewhere:

  • Website analytics supplies sessions and page activity.
  • Marketing automation supplies campaign and email engagement.
  • Event platforms supply registrations and attendance.
  • The CRM supplies lifecycle, ownership, opportunity, and sales outcomes.
  • Product analytics supplies signups, logins, and feature activity.
  • Form and chat systems supply direct requests and conversations.

Clay should not be described as natively collecting every possible behavioral signal. The relevant data points must be available in the table before a formula can evaluate them.

Score it only after an identity-resolution process has associated the activity with the appropriate record at an acceptable confidence level.

Normalize values before scoring

Raw data is often too inconsistent for dependable formulas. Normalize at least:

  • Job titles, functions, and seniority
  • Industry and subindustry names
  • Employee and revenue bands
  • Technology names
  • Countries, regions, and territories
  • Email-validity statuses
  • Null, unknown, and not-applicable values

For example, “VP Sales,” “Vice President of Sales,” and “Sales VP” may map to the same normalized function and seniority. Keep the raw title so operators can inspect the classification.

Handle nulls explicitly. A blank value should not accidentally match a fallback category or receive the same penalty as a confirmed disqualifier.

Record source, freshness, and confidence

Important attributes should have companion metadata:

  • Source: CRM, submitted form, enrichment provider, analytics platform, AI classification, or another origin
  • Last verified: When the value was observed or confirmed
  • Confidence: High, medium, low, or a controlled numerical scale
  • Conflict status: Whether another available source disagrees
  • Verification status: Whether a person or automated process reviewed it

A submitted role may be current but self-reported. An enriched title may be structured but stale. An AI-generated buying-role classification may help with triage but should not silently receive the same certainty as a verified CRM value.

The scoring response can remain simple: award full points to high-confidence values, withhold or reduce points for uncertain values, and route strategically important conflicts to review.

Define missing and contradictory data policies

For a missing value, choose one treatment:

  1. Award no points for that attribute.
  2. Mark the record incomplete and score the available fields.
  3. Trigger additional enrichment or manual research.

For contradictory values:

  1. Preserve each raw source value.
  2. Mark the normalized field as conflicting.
  3. Withhold or reduce the affected points.
  4. Verify high-value records before automated routing.
  5. Record which source ultimately won and why.

Resolve duplicates before account aggregation. Otherwise, the model may count the same person or event more than once, inflate account engagement, or route duplicate records to different owners.

Refresh according to signal volatility

Use separate refresh policies rather than one schedule for every field.

Industry, headquarters region, and broad employee band may change slowly. Job status, active opportunities, pricing activity, product use, and email responses can change faster.

Choose the cadence according to:

  • How quickly the signal changes
  • How quickly sales acts
  • Sales-cycle length
  • Enrichment cost
  • Harm caused by stale data
  • Source reliability

There is no universal daily or weekly schedule. Refresh valuable, volatile fields often enough to support their associated actions.

Privacy, consent, data licensing, retention, security, and regional restrictions should also be included in implementation review. Qualified internal privacy, security, and legal stakeholders should review the specific sources, matching methods, retention rules, and automated actions before launch.

4. Build the scoring table and Formula columns in Clay

Clay lists three prerequisites for custom lead scoring: available scoring data points, defined criteria, and CRM fields intended to receive synchronized scores. Its documented setup path is to add a table column, select Formula, enter criteria through the Formula Generator, review the result, and save it.

The broader architecture below is an implementation recommendation, not an official Clay-prescribed data model.

Recommended implementation sequence

  1. Ingest records. Include stable person, account, or CRM identifiers.
  2. Preserve raw values. Keep source fields unchanged for auditing.
  3. Normalize scoring fields. Convert inconsistent values into controlled categories.
  4. Enrich selectively. Request missing information only when it can affect the decision.
  5. Add quality metadata. Store source, freshness, confidence, and conflict status.
  6. Create component scores. Calculate fit, engagement, and penalties separately.
  7. Calculate the final score. Combine components without deleting them.
  8. Assign a tier. Apply boundaries, minimum requirements, and review rules.
  9. Generate a score reason. List the decisive positive and negative conditions.
  10. Review the distribution. Inspect the volume and characteristics of each tier.
  11. Test CRM writeback. Confirm field types, ownership, updates, and timestamps.
  12. Activate routing gradually. Enable production actions only after review.

Recommended table structure

A scoring table might contain:

  • Clay row identifier
  • CRM contact, lead, or account ID
  • Email and company domain
  • Raw title, industry, size, region, and technology fields
  • Normalized person and account attributes
  • Engagement fields, event timestamps, and rolling totals
  • Source and last-verified fields
  • Confidence and conflict fields
  • Fit Score
  • Engagement Score
  • Negative Score
  • Final Score
  • Tier
  • Score Reason
  • Scored At
  • Model Version
  • Next Action

The identifiers depend on whether the table is centered on people, companies, or an already-resolved combination of both.

Choose the appropriate output pattern

Clay documents three Formula output patterns: number-based scores, conditional grades, and binary Yes or No fit labels. The setup and output types are described in the official Formula-based scoring instructions.

These patterns can coexist. A table might contain a numeric Fit Score, a binary Minimum Fit field, and an operational Tier. The binary field can stop an engaged but ineligible record from qualifying through total points alone.

Keep raw attributes and component scores visible. If a record unexpectedly becomes Tier 1, an operator should be able to determine whether the cause was title normalization, an uncapped engagement count, a stale technology match, or an omitted penalty.

Make the reason inspectable

A useful Score Reason might read:

Target industry; 100–500 employees; director-level operations role; recent pricing-page visit; personal email requires verification.

The reason should list decisive inputs rather than restating “high score.” If AI generates or summarizes the explanation, validate it against the actual rule columns. An AI-written rationale should never replace the underlying conditions, values, and component scores.

Before scores control outreach or ownership, inspect:

  • High-scoring records
  • Records immediately above and below each boundary
  • Poor-fit records with high engagement
  • Strong-fit records with no engagement
  • Records with missing or conflicting data
  • Records affected by negative rules
  • Existing customers, competitors, and duplicates
  • Representative records from each region and segment

5. Translate the model into transparent scoring logic

Start by expressing the logic in plain language:

  • Fit Score is the sum of eligible firmographic, persona, and technographic points.
  • Engagement Score is the sum of eligible recent behavioral and timing points.
  • Negative Score is the nonnegative magnitude of applicable penalties.
  • Final Score equals Fit Score plus Engagement Score minus Negative Score.
  • Tier applies calibrated boundaries, minimum requirements, and review rules.

All examples below are illustrative pseudocode, not tested or official Clay Formula syntax.

Numeric component pattern

fit_score =
  employee_band_points
  + industry_points
  + seniority_points
  + technology_points

engagement_score =
  pricing_visit_points
  + demo_points
  + recent_engagement_points
  + buying_group_activity_points

negative_score =
  competitor_penalty_magnitude
  + invalid_contact_penalty_magnitude
  + excluded_region_penalty_magnitude
  + inactivity_penalty_magnitude

final_score =
  fit_score
  + engagement_score
  - negative_score

Every penalty in this convention is stored as a positive magnitude. A competitor penalty of 50 is stored as 50, not -50, because the total Negative Score is subtracted once.

Preserving the components also permits different treatment for the same total. Fit 45 plus Engagement 10 may require a different action from Fit 15 plus Engagement 40.

Binary qualification pattern

minimum_fit =
  target_industry_is_true
  AND supported_region_is_true
  AND relevant_role_is_true
  AND disqualifier_is_false

qualified_for_sales =
  minimum_fit
  AND (
    final_score >= calibrated_threshold
    OR eligible_hard_trigger_is_true
  )

This prevents a high total from masking a critical incompatibility.

Tier pattern

if contradictory_high_value_data:
  tier = "Review"

else if minimum_fit AND final_score >= tier_1_boundary:
  tier = "Tier 1"

else if minimum_fit AND final_score >= tier_2_boundary:
  tier = "Tier 2"

else:
  tier = "Tier 3"

The organization must calibrate the boundaries. Tiers should correspond to actions rather than exist only as reporting categories.

Combine hard triggers with minimum fit

A direct demo request may deserve immediate attention, but it does not always justify automatic sales ownership. A student, competitor, unsupported account, or existing customer could submit the same form.

if recent_demo_request AND minimum_fit:
  next_action = "Route to sales"

else if recent_demo_request AND fit_is_missing_or_conflicting:
  next_action = "Immediate manual review"

else if recent_demo_request AND confirmed_poor_fit:
  next_action = "Specialized nurture or disqualification review"

This preserves responsiveness without treating every submission as equivalent.

Use caps, windows, and decay carefully

Optional model controls include:

  • Point caps: Limit the contribution of repetitive low-value activity.
  • Rolling windows: Count only events within a relevant recent period.
  • Decay: Reduce engagement after inactivity.
  • Frequency rules: Reward repeated substantive actions within a short period.
  • Category caps: Prevent one signal category from dominating the score.

These are design patterns, not documented Clay defaults. The appropriate cap or period depends on the sales cycle and observed outcomes.

email_click_points =
  min(eligible_recent_email_clicks * low_value_weight, click_cap)

pricing_visit_points =
  count(pricing_visits_within_selected_window) * pricing_weight

decayed_engagement =
  recent_engagement_points * selected_recency_factor

Passive actions such as opens and general page views should not accumulate unlimited points. Otherwise, automation, long browsing histories, or repetitive low-value activity can overwhelm stronger fit evidence.

Worked illustrative record

A published third-party example assigns:

  • 20 points for 100–500 employees
  • 10 points for Series B or later funding
  • 15 points for using a competing tool
  • 10 points for director-or-higher seniority
  • 10 points for a pricing-page visit
  • 5 points for a case-study download

The total is 70. The same example routes scores of 60 or more to sales, 30–59 to nurture, and below 30 to awareness. These are illustrative figures for that scenario, not official Clay settings or validated universal thresholds (Factors’ sample model).

Component Points
Employee band 20
Funding stage 10
Competing technology 15
Seniority 10
Pricing-page activity 10
Case-study download 5
Final total 70

An explainable implementation would also show whether the employee count and funding data are current, whether the competing technology is confirmed, and when the behavioral events occurred.

High engagement with weak fit

Consider another record:

  • Several repeat visits
  • Recent webinar attendance
  • A pricing-page visit
  • Senior title inferred with low confidence
  • Conflicting industry values
  • Unsupported region according to one source
  • No verified company identifier

The record may earn a high Engagement Score, but automatic routing would be premature. It should enter Review or an appropriate nurture path until the company, region, and role are verified.

High engagement can justify attention without proving qualification. That is the main reason not to let a combined total erase its components.

6. Turn scores into CRM fields, routing, and account-level priorities

In the proposed architecture, the CRM remains the operational system of record. Clay acts as the scoring, enrichment, and research layer.

Recommended CRM fields include:

  • Fit Score
  • Engagement Score
  • Negative Score
  • Final Score
  • Tier
  • Score Reason
  • Scored At
  • Model Version
  • Data Confidence
  • Next Action

The CRM may also need review status, routing result, disqualification reason, and source-workflow fields.

Map tiers to actions

Tier Example outcome
Tier 1 Assign to a named account executive
Tier 2 Send to an SDR pool or round robin
Tier 3 Place in nurture or a research backlog
Review Hold automated outreach and request verification
Disqualified Suppress standard sales routing while retaining the reason

The appropriate mapping depends on territory design, ownership rules, lifecycle definitions, and sales capacity.

Third-party Clay workflows describe possible actions such as creating or updating a CRM record, notifying a representative, creating a task, applying a marketing tag, starting a sequence, invoking a webhook, or entering nurture. Treat these as architectural possibilities rather than proof that every action is available in every plan or integration (third-party RevOps workflow examples).

Verify current product documentation for connector availability, authentication, field mapping, round-robin configuration, and workflow limits before deployment.

Send the rationale with the score

Routing should include the reason a record qualified. A task or alert is more useful when it says:

  • Target account and supported region
  • Relevant director-level contact
  • Competing technology detected
  • Recent demo request
  • No active opportunity found
  • Data verified on a stated date

This gives sales context and makes incorrect rules easier to challenge. A bare “Tier 1” label is harder to trust and improve.

Prevent synchronization loops

Define:

  • Which system owns each field
  • Which fields Clay may update
  • Which changes trigger recalculation
  • Whether unchanged values should be written again
  • A Scored At timestamp
  • A Model Version
  • A workflow-run identifier, where appropriate
  • Rules suppressing updates created by the scoring workflow itself

For example, the CRM might own lifecycle stage and sales disposition, while Clay owns the latest computed score and reason. Closed outcomes can return to the scoring dataset without allowing the score to overwrite a sales-owned outcome.

Test these rules with nonproduction records or a controlled segment.

Score accounts without indiscriminate summation

In multi-contact buying motions, an account score should reflect relevant people rather than every activity from every contact.

Classify contacts by buying role:

  • Decision-maker
  • Champion
  • Evaluator
  • User
  • Procurement participant
  • Irrelevant contact
  • Unknown role

Then choose an aggregation method:

  • Sum: Rewards additional relevant contacts but can inflate the result.
  • Average: Reduces volume effects but may dilute strong decision-maker activity.
  • Top contact: Uses the strongest relevant contact but can overlook buying-group depth.
  • Selected-contact rollup: Aggregates only verified buying-group members.
  • Hybrid: Combines the strongest contact with a capped multi-contact bonus.

No method is universally superior. Deduplicate people and events first, cap repetitive behavior, and prevent one active but irrelevant contact from dominating the account.

Before launch, test routing, alerts, suppression, nurture, ownership, and every fallback path—including records that should not route.

7. Set thresholds with evidence, then backtest the model

Clay does not prescribe ideal point values, fit-versus-engagement weighting, tier boundaries, or a sales-ready threshold. Those choices belong to the implementing organization.

Useful inputs include:

  • Current ICP criteria
  • Characteristics of top customers
  • Sales and customer interviews
  • Common disqualification reasons
  • Territory and coverage constraints
  • Buying roles observed in successful deals
  • Closed-won and closed-lost records
  • Sales-accepted and sales-rejected leads
  • Opportunity creation by source or segment

Start small when history is limited

A team with little clean historical data should begin with a small rule-based baseline:

  • A few essential account-fit rules
  • A few relevant persona rules
  • A limited set of meaningful engagement signals
  • Explicit disqualifiers
  • A review path for uncertainty

The goal is not to represent every possible buying signal. It is to produce a model that sales and operations can inspect, understand, and challenge.

A large predictive or AI-generated model may appear sophisticated while relying on incomplete outcomes, biased labels, or unstable fields. Complexity can make errors harder to detect.

Use outcomes to inform relative weights

When usable CRM history exists:

  1. Calculate the overall rate at which eligible records became sales-accepted opportunities.
  2. Calculate the same rate for records with each candidate attribute.
  3. Repeat for employee bands, roles, technologies, engagement types, and other segments.
  4. Use relative differences to inform initial weights.
  5. Check whether the pattern persists across periods, regions, and acquisition sources.

One scoring guide recommends comparing segment-level conversion rates with the overall baseline and evaluating the completed model on separate historical records (Kubaru’s lead-scoring guide).

Correlation is not proof of causation. An attribute may appear successful because of existing sales coverage, campaign targeting, or incomplete data. Combine quantitative analysis with sales and customer context.

Backtest on separate records

Where possible, divide historical data into:

  • A design set used to identify rules and initial weights
  • A test set used to evaluate the completed model

Apply the formula to the test records as if their outcomes were unknown. Review:

  • Record volume in each tier
  • Sales acceptance and disqualification by tier
  • Opportunity creation
  • Stage conversion
  • Win rate
  • Sales-cycle length
  • Revenue
  • Time to first touch
  • Common false-positive and false-negative patterns

Preview both the number and characteristics of records in each band. If an operationally unmanageable share of the database becomes Tier 1, the model needs different boundaries or stronger qualification rules.

Use precision and recall as diagnostic concepts

Precision asks: of the records routed as qualified, how many were genuinely useful to sales?

Recall asks: of the valuable buyers in the evaluated population, how many did the model capture?

The appropriate balance depends on strategy and capacity. A small enterprise team may prefer fewer, deeply researched accounts. A high-volume inbound motion may tolerate broader routing with qualification downstream.

Review false positives and false negatives separately

For false positives, investigate:

  • Overweighted passive engagement
  • Stale firmographic data
  • Incorrect role classification
  • Duplicate events
  • Missing disqualifiers
  • A boundary set too low

For false negatives, investigate:

  • Overbroad negative rules
  • Omitted buying roles
  • Missing data treated as negative
  • Engagement that decayed too quickly
  • Ignored offline or product signals
  • An outdated ICP

Feed closed-won, closed-lost, sales-accepted, sales-rejected, and disqualified outcomes back into evaluation. Revise weights, caps, windows, decay rules, and tier boundaries as the ICP, sales process, or observed behavior changes.

Do not promise a universal performance lift or impose a fixed review interval. Review when enough new outcome data exists to learn from—and sooner when routing volume, product, territory, or funnel behavior changes materially.

8. Launch with governance, monitoring, and failure controls

A scoring workflow is not finished when the formula runs. It needs ownership, versioning, monitoring, and rollback controls.

Assign ownership and change authority

Name an accountable model owner and document who can:

  • Propose a new signal
  • Change normalization logic
  • Approve weights and thresholds
  • Modify CRM fields
  • Test a new version
  • Authorize deployment
  • Roll back a failed release
  • Resolve disagreements between sales and marketing

Without ownership, scoring models accumulate undocumented exceptions and conflicting interpretations.

Version every model

Every written score should be traceable to the rules and boundaries that produced it. Store a Model Version with the CRM score and maintain a change log containing:

  • Change reason
  • Attributes affected
  • Old and new logic
  • Test results
  • Expected operational impact
  • Approval record
  • Launch date
  • Rollback plan

Versioning allows operators to compare outcomes before and after a change rather than treating all historical scores as equivalent.

Monitor data and workflow failures

Create monitoring views for:

  • Missing required fields
  • Stale records
  • Conflicting values
  • Low-confidence classifications
  • Duplicate contacts or accounts
  • Unexpected score distributions
  • Sudden changes in qualification volume
  • CRM writeback failures
  • Records stuck without an owner
  • Routes that never execute
  • Suppressed records still receiving outreach
  • Scores produced by obsolete versions

Monitor distributions as well as averages. A stable average can hide a formula error that pushes more records into both the highest and lowest tiers.

Apply human review where the cost of error is high

Require review for strategically valuable records with contradictory company, role, region, or ownership data. Also review significant actions based on AI-generated classifications or explanations.

AI can help categorize titles or summarize reasons, but its output should be checked against source fields and deterministic rules before it controls ownership, suppression, or outreach.

If predictive scoring is introduced, run it alongside the explainable rule-based model first. Compare recommendations and investigate disagreements. Historical models can reproduce bias contained in past coverage, outcome labels, or selected features.

Audit negative rules

Review suppressed and disqualified records separately, especially those affected by ambiguous rules such as:

  • Personal email use
  • Career-page activity
  • Inactivity
  • Unknown industry
  • Missing technology
  • Consultant status
  • Geography inferred from conflicting sources

Preserve the reason for every penalty so false negatives can be identified.

Control privacy, cost, and data use

Before activating enriched or behavioral data, have appropriate internal and legal stakeholders review privacy, consent, licensing, retention, security, and regional requirements. Document which data is used, for which decision, how long it is retained, and who can access it.

Monitor enrichment and refresh costs as part of model performance. Do not enrich every possible field merely because it is available. A field should justify its cost by materially improving qualification, routing, or research.

Prelaunch checklist

Before scores control production workflows:

  • [ ] Validate every source field.
  • [ ] Preserve raw and normalized values.
  • [ ] Confirm null and conflict handling.
  • [ ] Inspect representative records manually.
  • [ ] Review records around every threshold.
  • [ ] Confirm explanations match the actual rules.
  • [ ] Test CRM field types and writeback.
  • [ ] Test sales, nurture, review, suppression, and fallback routes.
  • [ ] Confirm field ownership between Clay and the CRM.
  • [ ] Verify timestamps and model versions.
  • [ ] Check duplicate and account-rollup behavior.
  • [ ] Document monitoring and escalation.
  • [ ] Test rollback.
  • [ ] Obtain sales sign-off on tiers and handoffs.

Frequently asked questions

What buyer-scoring attributes does Clay officially document?

Clay documents examples such as employee count, company size, industry, job title, revenue, and technology stack. For engagement, it lists email responses, website visits, online behavior, and event participation.

These are examples rather than a required field list. Users must decide which attributes matter, define the criteria, and make the corresponding data points available.

Can Clay create numeric scores, letter grades, and binary fit labels?

Yes. Clay documents Formula columns for number-based point scores, conditional letter grades, and binary Yes or No fit labels.

These outputs can be combined. A workflow can retain numeric Fit and Engagement scores, calculate an operational grade, and use a binary minimum-fit rule to stop poor-fit records from qualifying through activity alone.

Should fit and engagement be separate Clay scores?

Usually. Fit describes whether a person or account matches the target market; engagement describes observed recent activity. They change at different rates and can imply different next actions.

A combined score can support prioritization, but the component scores should remain visible. This lets operators distinguish an excellent-fit but inactive account from a highly active record with weak or uncertain fit.

What score should trigger sales routing?

There is no universal Clay threshold. The correct boundary depends on the ICP, score range, acquisition channel, sales capacity, data quality, and cost of false positives and false negatives.

Choose an initial threshold from historical outcomes and sales input, preview the records it would route, and test it before automation. A hard trigger such as a demo request can accelerate review, but minimum-fit and disqualification rules may still apply.

Can Clay scores be written to HubSpot or Salesforce and used for routing?

A proposed architecture can write scoring outputs to CRM fields and use those fields in downstream assignment, task, alert, nurture, or suppression workflows. Third-party implementation sources describe workflows involving HubSpot and Salesforce, but exact connector setup, supported actions, field mapping, and routing configuration should be verified in current product documentation.

Use the CRM as the operational system of record, define field ownership, include Scored At and Model Version fields, and test for synchronization loops before enabling production routing.

Conclusion

Launch in stages. Start with a small, transparent set of fit, engagement, and disqualification rules. Implement them in visible Clay Formula columns, inspect the resulting distribution, and connect only validated tiers to CRM actions.

Preserve the data source, score reason, confidence, freshness, component scores, and model version so every result can be audited. Most importantly, treat every weight and threshold as a testable hypothesis. Improve the model with sales acceptance, disqualification, opportunity, win, and revenue outcomes rather than copying a third party’s example.