robots.txt, security.txt and Sitemaps in PagibleAI

The PagibleAI theme package serves /robots.txt, /security.txt and XML sitemaps dynamically. Editors maintain the robots.txt rules and the security contact data as configuration elements on the root page, and the sitemaps are generated from the published page tree.

This page explains how the routes work, which pages end up in the sitemaps and how to fix the most common problems with crawlers. Source: routes, RobotsController, SecurityController and SitemapController.

Overview of the generated URLs

With the default sitemap prefix, the theme package registers these public routes:

  • /robots.txt – published robots.txt rules plus the sitemap locations
  • /security.txt – RFC 9116 security contact information
  • /sitemap.xml – all indexable pages, or a sitemap index for very large sites
  • /sitemap-1.xml, /sitemap-2.xml, … – sitemap chunks with up to 50,000 URLs each
  • /sitemap-news.xml – Google News sitemap for recent news pages

The routes are registered outside the catch-all page route, so a prefix in CMS_PAGEROUTE does not change them. With multi-domain routing enabled (cms.multidomain), each route is bound to the requested domain and only returns data of that domain.

robots.txt

Remove the static robots.txt file

Fresh Laravel applications contain a static public/robots.txt file. Web servers, including php artisan serve, deliver that file before Laravel receives the request, so the dynamic route is never reached.

If the static file contains custom rules, first copy them into the robots.txt configuration element of the root page and publish the page. Then remove the file:

rm public/robots.txt

Alternatively, configure the web server to pass /robots.txt to Laravel even when the static file exists.

Add rules in the admin panel

  1. Open the root page of the site in the admin panel.
  2. Switch to the Config tab and add a new configuration element.
  3. Choose robots.txt from the expert group.
  4. Enter the rules in the text field, for example the lines below.
  5. Save and publish the root page.

Only the published configuration is used, so a saved draft has no effect until it is published. The field accepts up to 500,000 characters.

User-agent: *
Disallow: /search

User-agent: GPTBot
Disallow: /

Which root page is used

PagibleAI reads the robots.txt element only from root pages (pages without a parent) that are visible or hidden in navigation (status 1 or 2). It walks them in page-tree order and uses the first one with valid robots.txt text. Inactive root pages are skipped.

With multi-domain routing, only root pages assigned to the requested domain are considered, so every domain can have its own rules. Without multi-domain routing, the first matching root page of the whole tree wins.

Generated sitemap lines

The response contains the published rules first, followed by an empty line and Sitemap: declarations for the main sitemap and the news sitemap, using the configured sitemap prefix and the absolute URL of the current domain:

User-agent: *
Disallow: /search

Sitemap: https://example.com/sitemap.xml
Sitemap: https://example.com/sitemap-news.xml

If your rules already contain a Sitemap: line with one of these exact URLs (the keyword is case-insensitive), PagibleAI does not add that line a second time. Without any configured rules, /robots.txt contains only the two Sitemap: lines.

Limits, validation and caching

  • A leading UTF-8 byte order mark is removed and Windows or old Mac line endings are converted to \n.
  • Text with invalid UTF-8, control characters other than tab and line breaks, or only whitespace is ignored; PagibleAI then continues with the next root page.
  • The public response is limited to 500 KiB, the size common crawlers support. The editor limit counts characters, the controller counts encoded bytes. If the rules plus sitemap lines exceed the limit, only the sitemap lines are returned.
  • Invalid or oversized rules never break the route: the sitemap declarations are always served.
  • Responses are sent as text/plain; charset=utf-8 with Cache-Control: public, max-age=300 and an ETag, so crawlers and proxies can use conditional requests. Changes may take up to five minutes to appear in shared caches.

security.txt

/security.txt publishes the contact details for reporting vulnerabilities in the RFC 9116 format. It is generated from the security configuration element of the root page and uses the same root-page lookup as robots.txt: the first visible or hidden root page (of the requested domain with multi-domain routing) that has a security element. Remember to publish the root page after changing it.

Configure the security element

Open the root page, switch to the Config tab and add the security element from the expert group. It offers these fields:

  • Contact URL (contact, required) – absolute HTTPS link to contact the security team, e.g. a form
  • Contact email (email) – e-mail address for security reports, max. 254 characters; output as a second Contact: mailto: line
  • Expires (expires, required) – date after which the information is outdated; output as the end of that day in UTC, e.g. 2027-06-30T23:59:59Z
  • Encryption (encryption) – HTTPS link to the public key for encrypted messages
  • Acknowledgments (acknowledgments) – HTTPS link to the page thanking security researchers
  • Preferred languages (preferred-languages) – comma-separated language codes, e.g. en, de
  • Canonical (canonical) – HTTPS URL where the security.txt file is published
  • Policy (policy) – HTTPS link to the security and disclosure policy
  • Hiring (hiring) – HTTPS link to security-related jobs
  • CSAF (csaf) – HTTPS link to the CSAF provider metadata

All URL fields must be absolute https:// URLs with a host and at most 2,048 characters. Optional values that fail validation are silently left out of the response.

Example output

The fields are always written in this order and the response is cached for five minutes (Cache-Control: public, max-age=300):

Contact: https://example.com/security
Contact: mailto:security@example.com
Expires: 2027-06-30T23:59:59Z
Encryption: https://example.com/pgp-key.txt
Acknowledgments: https://example.com/hall-of-fame
Preferred-Languages: en, de
Canonical: https://example.com/.well-known/security.txt
Policy: https://example.com/disclosure-policy
Hiring: https://example.com/jobs
CSAF: https://example.com/.well-known/csaf/provider-metadata.json

If no security element exists, or if the contact URL or the expiry date is missing or invalid, /security.txt returns 404 Not Found. The expiry date is not compared with the current date, so renew it before it passes; RFC 9116 recommends an expiry less than a year in the future.

Serve /.well-known/security.txt

PagibleAI registers only /security.txt, not the /.well-known/security.txt location defined by RFC 9116. RFC 9116 allows redirects, so add one to routes/web.php of your application or configure an equivalent rewrite in the web server:

use Illuminate\Support\Facades\Route;

Route::redirect('/.well-known/security.txt', '/security.txt', 301);

XML sitemaps

Sitemap URLs and the CMS_SITEMAP prefix

The sitemap file names use the cms.theme.sitemap setting from config/cms/theme.php, which reads the CMS_SITEMAP environment variable and defaults to sitemap:

CMS_SITEMAP="sitemap"

With CMS_SITEMAP="pages", the URLs become /pages.xml, /pages-news.xml and /pages-1.xml. The Sitemap: lines in robots.txt follow automatically. Run php artisan config:clear (and php artisan route:cache again if you cache routes) after changing the value.

/sitemap.xml returns a single <urlset> with <loc> and <lastmod> (from the page's last update) as long as the site has at most 50,000 candidate pages. Above that, it returns a <sitemapindex> that points to /sitemap-1.xml, /sitemap-2.xml and so on, each with up to 50,000 URLs ordered by page ID. Chunk numbers outside the available range return 404.

Which pages are listed

Sitemaps only contain published data that every anonymous visitor can see:

  • Included: pages with status visible (1) and hidden in navigation (2)
  • Excluded: inactive pages (status 0)
  • Excluded: redirect pages, i.e. pages with a non-empty To URL
  • Excluded: pages with their own frontend access restriction (access roles), even if the requesting user could see them
  • Excluded: pages whose robots meta element is set to No index
  • Excluded: deleted pages and pages of other tenants
  • With multi-domain routing, only pages of the requested domain

Use hidden in navigation for landing pages that should be indexed but not appear in menus, and set the robots meta element to No index to keep a visible page out of the sitemap.

News sitemap

/sitemap-news.xml is a Google News sitemap. It applies the same exclusions as the main sitemap and additionally lists only pages that

  • use the page type news (offered by news-oriented themes such as Journal and News; the base theme only has page, docs and blog), and
  • were created within the last two days.

The newest 1,000 articles are included, sorted by creation date. Each entry contains the page title, the creation date as publication date, the page language and the publication name. The publication name is the Title of the website configuration element: PagibleAI walks the visible and hidden root pages (of the requested domain with multi-domain routing) in page-tree order and uses the first one whose website element has a title, so set it when you submit the news sitemap. Without matching pages, the file is a valid but empty <urlset>.

A database index on tenant_id, type, deleted_at and created_at added by the core migrations keeps this query fast; run php artisan migrate after upgrading.

Rate limiting and caching

All sitemap routes share the cms-sitemap rate limiter, which allows 10 requests per minute per IP address. Additional requests from the same IP, including requests for other sitemap files, receive HTTP 429 until the window expires. Sitemap responses are sent with Cache-Control: public, max-age=300, so a reverse proxy or CDN in front of the application can serve repeated requests without reaching Laravel.

To change the limit, for example for a site with many sitemap chunks, redefine the limiter in the boot() method of a service provider that runs after the PagibleAI theme provider, for example App\Providers\AppServiceProvider:

use Illuminate\Cache\RateLimiting\Limit;
use Illuminate\Support\Facades\RateLimiter;

RateLimiter::for('cms-sitemap', fn($request) =>
    Limit::perMinute(60)->by($request->ip())
);

Submit the sitemap to Google

Crawlers discover both sitemaps through /robots.txt. To get indexing reports, open Google Search Console, select the property for your domain, open Sitemaps and submit sitemap.xml (and sitemap-news.xml for news sites). For multi-domain setups, submit the sitemap of each domain in its own property. See Configure Google Search Console to also show Search Console data in the PagibleAI page metrics.

Troubleshoot robots.txt, security.txt and sitemaps

  • robots.txt shows the old Laravel content: the static public/robots.txt still exists. Move its rules to the root page, publish, and delete the file or let the web server pass the request to Laravel.
  • robots.txt contains only the sitemap lines: the root page with the rules is not published, is inactive, belongs to another domain, or the text contains invalid characters or exceeds 500 KiB.
  • robots.txt changes appear late: responses are publicly cacheable for five minutes; purge the CDN or proxy cache after urgent changes.
  • /security.txt returns 404: check that the root page is published and its security element has an HTTPS contact URL and an expiry date. http:// URLs are rejected.
  • An optional security.txt field is missing: the value failed validation, e.g. a non-HTTPS URL, an invalid e-mail address or language codes with unsupported characters.
  • A page is missing from the sitemap: check that it is published, not inactive, not a redirect, has no access roles and is not set to No index.
  • The news sitemap is empty: only pages of type news created within the last 48 hours are listed. Check the page type and creation date.
  • The news sitemap has no publication name: no visible or hidden root page of the domain has a published website element with a title.
  • Search engines get HTTP 429: the sitemap rate limiter allows 10 requests per minute per IP. Cache the sitemaps in a CDN or relax the limiter as shown above.
  • Sitemap or robots.txt URLs use the wrong host or scheme: Laravel builds the absolute URLs from the incoming request (and the requested domain with multi-domain routing). Configure trusted proxies so Laravel sees the public scheme and host.
  • The custom sitemap prefix is ignored: clear the configuration cache and rebuild a cached route list after changing CMS_SITEMAP.