Finding and fixing duplicate content
Duplicate content can lead Google to show the wrong page or rank none properly. Its technical causes can usually be fixed centrally.
What duplicates cost
Penalties are rare, but search engines pick the version they show. Link signals split across several URLs, and crawl resources go to waste.
Typical causes
URL variants with and without www, trailing slash or capitals
Parameters from filters, sorting or tracking
Print versions, pagination and session IDs
Product variants with near-identical text
Manufacturer descriptions reused across many shops
Publicly accessible staging or test environments
Finding duplicates
A crawl with Screaming Frog compares titles, headings and content. Google Search Console shows which pages Google treats as duplicates and which canonical URL it has chosen.
Fixing the causes
Technical duplicates are fixed at source with 301 redirects to one URL format and correctly set canonical tags. Parameter pages without search value are excluded from indexing, and test environments sit behind a password. Duplicated copy is merged or rewritten.
Approach
Causes are ranked by impact and fixed centrally in the CMS or server configuration. That stops new duplicates from appearing.