Pillar guide

HubSpot Data Cleanup: A Revenue-First Playbook

Most HubSpot clean-up projects fail because they chase tidiness instead of money. This guide gives you a triage framework, the ten highest-value jobs, safe-change rules, and a 30/60/90 sequence that fixes the defects actually costing you revenue — without a single bulk delete.

14 minute read · written by knovaly

Why most clean-up projects fail

Ask a RevOps lead when they last ran a HubSpot clean-up and most will describe the same pattern: a project kicked off with good intentions, a spreadsheet of everything wrong with the portal, three weeks of effort, and then silence. Six months later the same duplicates are back, the same deals are missing amounts, and nobody can say what the clean-up actually achieved. The project did not fail because the team lacked discipline. It failed because it optimised for tidiness rather than revenue.

A tidiness-first clean-up treats every defect as equally urgent. It standardises picklists, renames properties, deduplicates every contact regardless of deal value, and produces a portal that looks better in a property inspector but has not moved a single commercial number. Nobody in the leadership team notices, because none of the fixes touched pipeline coverage, forecast accuracy, or a renewal that was about to be missed. The project loses executive sponsorship at the first budget review, and the backlog reappears.

A revenue-first clean-up starts from the opposite question: which data defects, if fixed this week, change a number a VP of Sales or a CFO actually looks at? That reframing changes everything about sequencing, scope and who gets asked to help — and it is the only version of a clean-up project that survives contact with a busy sales team.

Clean in the order that unlocks money

Not all HubSpot fields carry equal commercial weight. A missing hs_lead_status on a marketing-qualified contact is a nuisance. A missing amount on an open deal at Contract Sent stage is a forecast error that could be worth six figures. Revenue-first sequencing means working outward from the properties that directly feed pipeline value, forecast category and renewal timing, and only then moving on to hygiene that supports segmentation and reporting quality.

A simple sequencing test

Before adding any defect to the clean-up backlog, ask three questions: does this field feed a forecast, a pipeline report, or a renewal date used anywhere in the business? Does fixing it change which deals a rep works this week? Would a board member notice the difference in a report? If the honest answer to all three is no, the defect belongs in the "fix by process" or "leave alone" bucket described in the next section, not at the top of week one.

The triage framework

Every defect uncovered during a clean-up audit should be sorted into exactly one of three buckets before anyone starts fixing anything. This single discipline is what stops a clean-up project sprawling into an open-ended tidiness exercise.

  • Fix now — high revenue impact, low effort to fix, and the fix does not depend on a policy decision. Missing amount on an open deal, an unowned deal above your average deal size, or a closedate set years in the past all belong here.
  • Fix by process — real impact, but the fix needs a rule change, a workflow, or a required property to stop it recurring, rather than a one-off manual correction. Inconsistent pipeline stages and unmanaged imports usually sit here.
  • Leave alone — low volume, low value, or genuinely ambiguous (a closed-lost deal from three years ago with a blank industry field is not worth anyone's time).
Triage decision table
DefectRevenue impactEffort to fixBucket
Open deal, missing amount, stage ≥ ProposalHighLowFix now
Open deal, closedate in the pastHighLowFix now
Deal with no hubspot_owner_idHighLowFix now
Duplicate companies with active dealsMedium–HighMediumFix now
Inconsistent dealstage usage across teamsMediumHigh (needs agreement)Fix by process
Unmanaged CSV imports creating new duplicates weeklyMediumHighFix by process
Contact missing lifecyclestage, no deal associationLow–MediumMediumFix by process
Closed-lost deal from 2021 with a blank industry fieldNegligibleLowLeave alone

The ten highest-value clean-up jobs

These ten categories account for the majority of recoverable value knovaly typically finds in a portal's first read-only scan. Each one is described with how to find it and how to fix it without creating a new mess.

1. Deals missing amount

Find it: build a deal view filtered to hs_is_closed = false and amount is unknown, sorted by dealstage descending. Fix it: require reps to enter a realistic amount before a deal can progress past your second pipeline stage — not at creation, where deal size is often genuinely unknown. Where deals are stuck with no amount and no recent activity, treat them as candidates for closing lost rather than guessing a number.

2. Deals missing or past closedate

Find it: filter open deals where closedate is blank or earlier than today. Fix it: for blanks, ask the owning rep for a realistic date this week; for past dates, either update to a genuine next milestone or move the deal to closed lost. Do not mass-update every past closedate to "today plus 30 days" — that manufactures a forecast spike that will embarrass whoever presents it.

3. Unowned records

Find it: filter deals and companies where hubspot_owner_id is empty, cross-referenced against active pipeline stages. Fix it: assign by territory or account list rather than round-robin for anything with existing deal value; round-robin is fine for genuinely new, unworked records.

4. Duplicate companies and contacts

Find it: HubSpot's native duplicate management tool, supplemented by a domain-based check for companies HubSpot's fuzzy match misses. Fix it: merge, do not delete — merging preserves associations, engagement history and deal links. Merge in batches of 50 or fewer and prioritise duplicates attached to open deals first.

5. Orphaned deals with no associations

Find it: a deal view showing deals with no associated contact and no associated company — these usually come from manual creation or a broken import mapping. Fix it: associate manually where the deal has real value; archive (never delete outright) where it is a genuine test record or import error, keeping a log of what was archived and why.

6. Contacts with no lifecyclestage

Find it: a contact list filtered to lifecyclestage is unknown, segmented by whether the contact has an associated deal or recent activity via notes_last_contacted. Fix it: backfill using deal association first (any contact on an open or won deal should never be blank), then apply a workflow to set lifecyclestage automatically going forward based on form submission or deal creation.

7. Stale open deals

Find it: open deals where hs_last_sales_activity_timestamp is older than your average sales cycle length. Fix it: route these to owners for a status review with a hard deadline — reclassify as active, close lost, or explicitly mark as "paused, review in Q3" rather than leaving them in limbo indefinitely inflating pipeline coverage.

8. Inconsistent pipeline stages

Find it: compare stage usage across teams or pipelines — if one team's "Proposal Sent" means the same thing as another's "Negotiation," your stage-based forecasting is comparing unlike things. Fix it: this is a process fix, not a data fix — get pipeline owners to agree stage definitions in writing, then remap historical deals in a single reviewed batch rather than leaving old and new definitions to coexist.

9. Bad country and currency data

Find it: free-text country fields with inconsistent spelling ("USA," "United States," "U.S.") and deals recorded in the wrong currency for the associated company's region. Fix it: convert country to a HubSpot dropdown property with a fixed list of values, then run a one-time normalisation pass; verify currency against company records before any bulk currency correction, since this directly affects reported deal value.

10. Unmanaged imports

Find it: review import history under Data Management for imports with no clear owner, no deduplication key, or an unusually high record count relative to what was expected. Fix it: require a named owner and a deduplication key (email, domain, or external ID) on every future import, and audit the last 12 months of import history for the ones most likely to have seeded today's duplicates.

Safe-change rules for bulk edits

The fastest way to turn a clean-up project into an incident is to make irreversible changes at scale before you have confirmed the fix is correct. These rules apply to every bulk action described above.

  1. Export before you touch anything. A full CSV export of the object list, with all relevant properties, takes minutes and gives you a rollback path no HubSpot undo feature can match.
  2. Work in batches of 50–100. Bulk edits across thousands of records at once make it impossible to spot an error until it has already propagated everywhere.
  3. Never bulk-delete. Archive instead — archived records can be restored, deleted ones generally cannot, and archived records still show up correctly in historical reporting.
  4. Prefer reassignment over deletion for ownership gaps. An unowned deal is a routing problem, not evidence the deal should not exist.
  5. Log every bulk edit. Record what changed, the filter used to select records, the date, and who approved it — a simple shared document is sufficient, but it must exist before the edit, not after someone asks what happened.
  6. Test on a small segment first. Run any new bulk edit or workflow against 10–20 records and manually verify the outcome before scaling it to the full list.

Stopping defects recurring

A clean-up that is not followed by prevention is a clean-up you will repeat every quarter. Two mechanisms do most of the work: workflows that catch defects as they are created, and required properties applied narrowly enough not to backfire.

Workflows as a first line of defence

A workflow that enrols any deal created without an owner and immediately notifies a manager catches the unowned-record problem within minutes rather than at the next audit. Similarly, a workflow that flags any deal moving into a late pipeline stage with a blank amount stops the highest-value defect from ever reaching your forecast. These workflows are cheap to build and, unlike a required property, they can notify a human rather than silently blocking a rep's work.

Where required fields backfire

Required properties feel like the obvious solution, but applied broadly they train reps to enter placeholder values just to move past a form — a dummy amount of £1, a closedate a decade in the future, a lifecyclestage set incorrectly because the form would not submit otherwise. These placeholder entries are worse than blank fields because they look valid in a report and quietly corrupt averages. The fix is to make properties required only at the pipeline stage where the information is genuinely knowable, and to pair them with validation rules (a closedate must fall within the next 12 months, an amount must be greater than zero) rather than a blanket "must not be blank."

Before and after: measurable outcomes

A clean-up project justifies itself with numbers a leadership team can see, not a subjective sense that the portal "feels" better. Track these three before starting and again 90 days after.

Typical before/after outcomes from a revenue-first clean-up
MetricBeforeAfter 90 days
Forecast accuracy (committed vs. actual closed-won)±35–45%±10–15%
Percentage of open deals with valid amount and closedate55–65%90%+
Time to build a trusted pipeline reportHalf a day, with manual filtering of bad dataMinutes, filters apply cleanly
Duplicate company rate8–12% of company recordsUnder 2%
Report trust (leadership willing to present a HubSpot number unchecked)Rare — numbers verified manually firstStandard practice

A 30/60/90 clean-up sequence

This sequence assumes a mid-sized portal and a small cross-functional team: a RevOps or sales-ops owner running the mechanics, and sales management making the judgement calls on deal status and ownership.

Days 1–30: fix now

Owner: RevOps, with sales management sign-off on deal-status decisions. Run the audit, apply the triage framework, and clear the entire fix-now bucket — missing amounts on late-stage deals, past closedates, unowned high-value records, and duplicates attached to open pipeline. Export a backup before every batch. Target: every open deal above your average deal size has a valid owner, amount and closedate.

Days 31–60: fix by process

Owner: RevOps, with pipeline-stage definitions agreed by sales leadership. Build the workflows and validation rules that stop the fix-now defects recurring. Reconcile inconsistent pipeline stages across teams. Standardise country and currency fields. Establish a named owner and deduplication key requirement for all future imports.

Days 61–90: institutionalise and measure

Owner: RevOps owns the cadence, sales management owns the weekly review. Set up standing active lists for each of the ten defect categories with a weekly or monthly review cadence attached. Calculate a baseline CRM health score and present the before/after numbers from the section above to leadership. Schedule the next full re-scan for 90 days out.

Measuring progress with a health score

"The data feels cleaner" is not a metric anyone can act on. A composite CRM health score — built from the percentage of open deals with a valid amount and closedate, the duplicate rate across companies and contacts, the percentage of records with a valid owner, and the proportion of contacts with a correctly set lifecyclestage — gives you a single number that moves in a predictable direction as you clear the triage buckets. Track it monthly, not quarterly; monthly tracking catches regressions from a bad import or a paused workflow before they compound into another full clean-up project. knovaly calculates a 0–100 Opportunity Score alongside its revenue categories specifically so this tracking does not require a manual audit every time.

What to do about legacy records nobody owns

Every portal accumulates a pile of records nobody wants to claim: deals from a rep who left two years ago, companies imported during a merger that was never fully integrated, contacts from an event list with no clear next step. These records are rarely worth a dedicated project, but leaving them unaddressed skews every aggregate report that does not explicitly filter them out.

The pragmatic approach is a single reassignment pass, not an ongoing debate. Reassign departed-rep pipeline to the current territory owner or a sales manager acting as placeholder, with a 30-day deadline to review and reclassify. For legacy companies and contacts with no recent activity and no open deal, archive rather than delete, and record the reason and date in your bulk-edit log. If a genuine dispute exists over who should own a segment of legacy records, that is a sales-management decision to resolve explicitly, not something to leave unowned indefinitely — an unowned record is functionally the same as a lost one every time a report filters by owner.

Frequently asked questions

Should we clean up HubSpot before or after a big campaign or renewal push?
Before, but only the portion of the clean-up that touches the segment you are about to activate. If you are launching a renewal campaign, fix closedate and amount on renewal-stage deals and confirm owner assignment for that book of business first — that is a two-day job, not a two-month one. Full-portal clean-up can run in parallel without blocking the campaign. Sequencing it this way means the campaign benefits immediately instead of waiting behind an open-ended tidiness project.
Is it safe to bulk-delete duplicate contacts and companies in HubSpot?
Treat deletion as a last resort, not a first move. HubSpot's merge tools consolidate engagement history, associations and email logs in a way that a delete-and-recreate approach cannot replicate, and a wrongly deleted record with attached deal history is very hard to reconstruct. Export a full backup of the object list before merging anything, merge in batches of no more than 50, and keep a log of every merge ID. Never enable a workflow or import step that auto-deletes records.
How long does a proper HubSpot data clean-up actually take?
For a mid-sized portal (10,000–50,000 contacts, 2,000–5,000 deals), the ten highest-value jobs in this guide typically take a combined 15–25 hours of hands-on work spread over 4–6 weeks, plus ongoing governance. The constraint is rarely the mechanical fix — it is agreeing ownership, validating pipeline stage definitions with sales leadership, and getting reps to confirm which of their deals are genuinely still live. Portals with heavy import history or multiple legacy CRMs merged in take longer.
Do required properties in HubSpot actually solve the underlying problem?
Partially, and only if used narrowly. Making closedate or amount required on deal creation stops new bad data entering, but it does nothing for the backlog already in the system, and if applied too broadly it trains reps to enter placeholder values just to get past the form — a dummy amount of £1 or a closedate a decade out is worse than a blank field because it looks legitimate in a report. Pair required fields with validation rules and stage-specific requirements rather than blanket mandates.
What is the single highest-leverage clean-up job if we can only do one?
Fixing missing or past closedate on open deals. This one property distorts pipeline coverage ratios, forecast category rollups, and every stage-velocity report built on hs_lastmodifieddate deltas. A deal sitting at a closedate eight months in the past with no activity is either dead and should be closed lost, or alive and needs a realistic date — either correction immediately improves forecast accuracy without touching a single other object.
Can knovaly clean up our HubSpot data for us?
No — knovaly connects read-only and never edits or deletes records in your portal. What it does is show you, ranked by pound value, exactly which defects are costing you recoverable revenue right now: which open deals are stalling, which renewals are unmanaged, which dormant accounts still have money attached. You then action the fixes yourself, or hand the list to whoever owns HubSpot admin, using the safe-change rules in this guide.
How do we stop the same defects reappearing three months after clean-up?
Assign an owner to each of the ten defect categories, put a lightweight active list against each one in HubSpot (for example, 'Open deals, no amount, stage > Qualified'), and review those lists on a fixed cadence — weekly for pipeline hygiene, monthly for contact and lifecycle hygiene. Combine this with narrow required-property rules at the right pipeline stage and a quarterly re-scan of the whole portal. Clean-up without a maintenance cadence decays within one sales cycle.
Who should own HubSpot data quality — sales ops, RevOps, or sales management?
RevOps or sales operations should own the mechanics — the workflows, required properties, and reporting that surface defects. Sales management must own the judgement calls: which deals in the 'unowned' or 'stale' pile are actually reassignable, which deserve to be closed lost, and which pipeline stage definitions have drifted. Clean-up projects stall when ops teams try to make commercial calls alone, or when sales leaders are asked to do manual data entry instead of decisions.