The original article recommended making three numbers match: the crawl, sitemap and search index estimate. That rule was too simplistic.
Those sources describe different things. The useful work is finding actual URL problems and deciding what each page should do, not forcing totals to agree.
Build an inventory with provenance
Combine a crawl, the sitemap and Search Console evidence. Keep the source and date for each observation. A site: search is not a complete inventory, and changes in its estimated count do not diagnose duplicate removal.
For alternate hostnames, test redirects and canonical signals directly rather than expecting a particular count from separate searches.
Review page types before excluding them
Categories, tags, archives and filtered product pages can have different purposes. Some help customers find useful content; others may create redundant or unwanted URLs.
Inspect representative pages and their links before applying a template-wide rule. A temporary redirect is not inherently a quality problem either. Determine whether its purpose and implementation are appropriate.
Apply the right treatment
Keep useful pages, consolidate genuine duplicates and redirect moved content to relevant equivalents. Use indexing controls deliberately, and ensure a crawler can see any noindex directive you rely on.
Generate the sitemap from the intended indexable canonical pages. Don’t use the number of products or posts as a universal target for the whole site.
Let the evidence develop
Record changes and allow for recrawling and reporting delays. Avoid stacking unrelated edits merely because a ranking moved for a day; that makes diagnosis harder. There is no universal stabilisation pattern that tells you exactly when to make the next change.
See the unwanted-URL review workflow and our SEO audit checklist for a more practical starting point.
