Web Log Archive · Index · Part 1 · 2 · 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 · 16 · 17 · 18 · 19 · 20 · 21 · 22 · 23 · Expand · Web Feed

Common navigation links on dynamic pages can produce (partial) duplicated content (identical body text served from different URLs). To prevent search engine algos from filtering or even penalizing these URLs, eliminate the overlapping content.

Does Google systematically wipe out duplicated content? If so, does it affect partial dupes too? Will Google apply site-wide 'scraper penalties' when a particular dupe-threshold gets reached or exceeded?

Following many 'vanished page posts' with links on message boards and usenet groups, and monitoring sites I control, I've found that indeed there is kinda pattern. It seems that Google is actively wiping dupes out. Those get deleted or stay indexed as 'URL only', not moved to the supplemental index.

Example: I have a script listing all sorts of widgets pulled from a database, where users can choose how many items they want to see per page (values for #of widgets/page are hard coded and all linked), combined with prevŠnext-page links. This kind of dynamic navigation produces tons of partial dupes (content overlaps with other versions of the same page). Google has indexed way too many permutations of that poorly coded page, and foolishly I didn't take care of it. Recently I got alerted as Googlebot-Mozilla requested hundreds of versions of this page within a few hours. I've quickly changed the script, putting a robots NOINDEX meta tag when the content overlaps, but probably too late. Many of the formerly indexed (cached, appearing with title and snippets on the SERPs) URLs have vanished, respectively became URL-only listings. I expect that I'll lose a lot of 'unique' listings too, because I've changed the script in the middle of the crawl.

I'm posting this before I've solid data to backup a finding, because it is a pretty common scenario. This kind of navigation is used at online shops, article sites, forums, SERPs ... and it applies to aggregated syndicated content too.

I've asked Google whether they have a particular recommendation, but no answer yet. Here is my 'fix':

Define a straight path thru the dynamic content, where not a single displayed entry overlaps with another page. For example if your default value for items per page is 10, the straight path would be:
Then check the query string before you output the page. If it is part of the straight path, put an INDEX,FOLLOW robots meta tag, otherwise (e.g. start=16&items=15) put NOINDEX.

I don't know whether this method can help with shops using descriptions pulled from a vendor's data feed, but I doubt it. If Google can determine and suppress partial dupes within a site, it can do that with text snippets from other sites too. One question remains: how does Google identify the source?

Monday, August 15, 2005

Thoughts on Duplicate Content Issues with Search EnginesNext Page

Previous PageHow to Gain Trusted Connectivity

Web Log Archive · Index · Part 1 · 2 · 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 · 16 · 17 · 18 · 19 · 20 · 21 · 22 · 23 · Expand · Web Feed

Author: Sebastian
  Web Feed

· Home

· Internet

· Blog

· Web Links

· Link to us

· Contact

· What's new

· Site map

· Get Help

Most popular:

· Site Feeds

· Database Design Guide

· Google Sitemaps

· smartDataPump

· Spider Support

· How To Link Properly

Free Tools:

· Sitemap Validator

· Simple Sitemaps

· Spider Spoofer

· Ad & Click Tracking

Search Google
Web Site

Add to My Yahoo!
Syndicate our Content via RSS FeedSyndicate our Content via RSS Feed

To eliminate unwanted email from ALL sources use SpamArrest!


neat CMS:
Smart Web Publishing

Text Link Ads

Banners don't work anymore. Buy and sell targeted traffic via text links:
Monetize Your Website
Buy Relevant Traffic

[Editor's notes on
buying and selling links

Digg this · Add to del.icio.us · Add to Furl · We Can Help You!

Home · Categories · Articles & Tutorials · Syndicated News, Blogs & Knowledge Bases · Web Log Archives

Top of page

No Ads

Copyright © 2004, 2005 by Smart IT Consulting · Reprinting except quotes along with a link to this site is prohibited · Contact · Privacy