|
The robots.txt file is an ASCII text file that has specific instructions for search engine robots about specific content that they are not allowed to index. These instructions are the deciding factor of how a search engine indexes your website's pages. The universal address of the robots.txt file is: www.domain.com/robots.txt. This is the first file that a robot visits. It picks up instructions for indexing the site content and follows them. This file contains two text fields. Lets study this example:
User-agent: *
Disallow:
The User-agent field is for specifying robot name for which the access policy follows in the Disallow field. Disallow field specifies URLs which the specified robots have no access to. An example:
User-agent: *
Disallow: /
Here "*" means all robots and "/ " means all URLs. This is read as, “No access for any search engine to any URL" Since all URLs are preceded by "/ " so it bans access to all URLs when nothing follows after "/ ". If partial access has to be given, only the banned URL is specified in the Disallow field. Lets consider this example:
# Research access for Googlebot.
User-agent: Googlebot
Disallow:
User-agent: *
Disallow: /concepts/new/
Here we see that both the fields have been repeated. Multiple commands can be given for different user agents in different lines. The above commands mean that all user agents are banned access to /concepts/new/ except Googlebot which has full access. Characters following # are ignored up to the line termination as they are considered to be comments.
User-agent: *
Disallow: /stats/
However, it is easy for a snooper to guess what you are trying to hide and simply typing the URL www.domain.com/stats in his browser would enable access to the same. This calls for one of the following remedies -
1. Change file names:
Change the stats filename from index.php to something different, such as stats- new.php so that your stats URL now becomes www.domain.com/stats/stats-new.php
Place a simple text file containing the text, "Sorry you are not authorized to view this page", and save it as index.php in your /stats/directory.
This way the snooper cannot guess your actual filename and get to your banned content.
2. Use login passwords:
Password-protect the sensitive content listed in your robots.txt file.
Optimization of the robots.txt file
The Right Commands in robots.txt :
Use correct commands. Most common errors include - putting the command meant for "User-agent" field in the "Disallow field" and vice-versa.
Please also note that there is no "Allow" command in the standard robots.txt protocol. Content not blocked in the "Disallow" field is considered allowed. Currently, only two fields are recognized: "The User-agent field" and the "Disallow field". Experts are considering the addition of more robot recognizable commands to make the robots.txt file more Webmaster and robot friendly.
 |
Note: Google is the only search engine, which is experimenting with certain new robots.txt
|
| |
commands. It recognises the "allow" command. Please read more details on the google site for robots.txt usage. |
Bad Syntax:
Do not put multiple file URLs in one Disallow line in the robots.txt file. Use a new Disallow line for every directory that you want to block access to. Incorrect Robots.txt
Example:
User-agent: *
Disallow: /concepts/ /links/ /images/
Correct robots.txt example:
User-agent: *
Disallow: /concepts/
Disallow: /links/
Disallow: /images/
Files and Directories:
If a specific file has to be disallowed, end it with the file extension and without a forward slash in the end. Study the following robots.txt example:
For file:
User-agent: *
Disallow: /hilltop.phpl
For Directory:
User-agent: *
Disallow: /concepts/
Remember if you have to block access to all files in the directory, you don't have to specify each and every file in robots.txt. You can simply block the directory as shown above. Another common error is leaving out the slashes altogether. This would leave a very different message than intended.
The Right Location for the robots.txt file:
No robot will access a badly placed robots.txt file. Make sure that the location is www.domain.com/robots.txt.
Capitalization in robots.txt
Never capitalize your syntax commands. Directory and filenames are case sensitive in Unix platforms. The only capitals used per standard are: "User-agent " and "Disallow"
Correct Order for robots.txt : If you want to block access to all but one or more than one robot, then the specific ones should be mentioned first. Lets study this robots.txt example:
User-agent: *
Disallow: /
User-agent: MSNbot
Disallow:
In the above case, MSNbot would simply leave the site without indexing after reading the first command. Correct syntax is:
User-agent: MSNbot
Disallow:
User-agent: *
Disallow: /
The robots.txt file :
Not having a robots.txt file at all could generate a 404 error for search engine robots, which could redirect the robot to the default 404-error page or your customized 404-error page. If this happens seamlessly, it is up to the robot to decide if the target file is a robots.txt file or an html file. Typically it would not cause many problems but you may not want to risk it. It's always a better idea to put the standard robots.txt file in the root directory, than not having it at all.
The standard robots.txt file for allowing all robots to index all pages is:
User-agent: *
Disallow:
Using # Carefully in the robots.txt file:
Adding comments after the syntax commands is not a good idea using "#". Some robots might misinterpret the line although it is acceptable as per the robots exclusion standard. New lines are always preferred for comments.
Using the robots.txt file
- Robots are configured to read text. Too much graphic content could render your pages invisible to the search engine. Use the robots.txt file to block irrelevant and graphic-only content.
- Indiscriminate access to all files, it is believed, can dilute relevance to your site content after being indexed by robots. This could seriously affect your site's ranking with search engines. Use the robots.txt file to direct robots to content relevant to your site's theme by blocking the irrelevant files or directories.
- The robots.txt file can be used for multilingual websites to direct robots to relevant content for relevant topics for different languages. It ultimately helps the search engines to present relevant results for specific languages. It also helps the search engine in its advanced search options where language is a variable.
- Some robots could cause severe server loading problems by rapid firing too many requests at peak hours. This could affect your business. By excluding some robots that might be irrelevant to your site, in the robots.txt file, this problem can be taken care of. It is really not a good idea to let malevolent robots use up precious bandwidth to harvest your emails, images etc.
- Use the robots.txt file to block out folders with sensitive information, text content, demo areas or content yet to be approved by your editors before it goes live.
The robots.txt file is an effective tool to address certain issues regarding website ranking. Used in conjunction with other SEO strategies, it can significantly enhance a website's presence on the net.
Article last updated: 11th March 2004
Related Reading:
A Standard for Robots Exclusion.
Guide to The Robots Exclusion Protocol
W3C Recommendations
Read about robots meta tag
© Copyright 2005, RedAlkemi
------------------------------------------------------------------------------------------------------------
This Article is Copyright protected. If you have comments; or would like to have this article republished free on your site, please contact the author here: SEO Articles Feedback. We just require all due credits carried; and text, hyperlinks and headers unaltered. This article must not be used in unsolicited mail.
E-mail this page |