SEO
robots.txt Validator
ตรวจสอบ robots.txt — โครงสร้างกลุ่ม User-agent คำสั่ง และทดสอบพาธที่กำหนดว่าอนุญาต/บล็อก
หากระบุเฉพาะโดเมน ระบบจะเติม /robots.txt ให้ อาจใช้ไม่ได้กับทุกเว็บไซต์ — บางเว็บไซต์บล็อกคำขอข้ามโดเมน (CORS) หากโหลดไม่สำเร็จ ให้เปิด robots.txt ในแท็บใหม่แล้ววางเนื้อหาด้วยตนเอง
robots.txt controls which paths of a site search bots are allowed to crawl. A mistake in this file can accidentally block an entire site from indexing, or expose internal sections that shouldn't be crawled. This tool checks the file's structure and tests a specific path against allow/disallow rules.
How to use it
- Paste the contents of robots.txt and the tool parses the User-agent groups and directives (Allow, Disallow, Sitemap).
- Enter a specific path (e.g. /admin/ or /blog/post-1) to check whether it's allowed for a chosen bot.
- Syntax errors and suspicious constructs are flagged separately from valid directives.
Common uses
- Checking before a deploy that a new robots.txt hasn't accidentally blocked important pages.
- Debugging why a specific page isn't indexed — it may be caught by a Disallow rule.
- Comparing rules for different bots (Googlebot, Bingbot, etc.) when separate User-agent groups target them.
Things to keep in mind
robots.txt only blocks crawling — it doesn't guarantee a page won't appear in the index: if something else links to it, Google can still show the URL with no description. To reliably exclude a page, use a meta noindex tag, and the page must not be blocked in robots.txt for that tag to be seen.
Paths in robots.txt are case-sensitive — /Page and /page are treated as different paths.
บทความเกี่ยวกับเครื่องมือนี้: robots.txt: บอตค้นหาตัดสินใจอย่างไรว่าอะไรสามารถรวบรวมข้อมูลได้
คำถามที่พบบ่อย
ทำไมกฎ Disallow ของฉันถึงไม่บล็อกหน้าที่ฉันคาดหวัง?
กฎ robots.txt จะจับคู่ตามคำนำหน้า ไม่ใช่หน้าที่ตรงเป๊ะ และกฎที่เฉพาะเจาะจงกว่าจะแทนที่กฎทั่วไป — Disallow: /blog จะไม่บล็อก /blog-archive เว้นแต่คุณจะคำนึงว่าการจับคู่คำนำหน้าทำงานอย่างไรจริง ๆ และกฎ Allow ที่อื่นอาจแทนที่ Disallow ที่กว้างกว่าได้
robots.txt ป้องกันไม่ให้หน้าเว็บถูกจัดทำดัชนีได้จริงหรือไม่?
ไม่น่าเชื่อถือด้วยตัวมันเองเพียงอย่างเดียว มันเพียงขอให้ crawler ที่มีมารยาทไม่ดึงหน้านั้น — หน้าที่ถูกห้ามยังคงถูกจัดทำดัชนีได้ (โดยไม่มีเนื้อหา) หากหน้าอื่นลิงก์มาที่หน้านั้น ดังนั้นควรใช้ meta tag หรือ header noindex เพื่อการรับประกันที่แท้จริง
เนื้อหา robots.txt ของฉันถูกส่งไปที่ไหนเพื่อตรวจสอบหรือไม่?
ไม่ การตรวจสอบทำงานทั้งหมดในเบราว์เซอร์ของคุณ ไม่มีการส่งข้อมูลไปยังเซิร์ฟเวอร์
Googlebot รองรับไวล์การ์ด * และ $ หรือไม่?
รองรับ — Googlebot และ crawler สมัยใหม่ส่วนใหญ่รองรับ * (จับคู่กับสตริงอักขระใด ๆ) และ $ (ทำเครื่องหมายจุดสิ้นสุดของ URL) แต่นี่เป็นส่วนขยายที่เกินกว่าข้อกำหนดดั้งเดิมปี 1994 ดังนั้นไม่ใช่ทุกบอตที่จะรองรับ
พาธใน robots.txt แยกความแตกต่างตัวพิมพ์ใหญ่-เล็กหรือไม่?
ใช่ /Admin/ และ /admin/ ถือเป็นสองพาธที่แตกต่างกัน ดังนั้นกฎ Disallow ที่เขียนตัวพิมพ์ผิดจะไม่มีผลกับพาธจริงเลย