Departure Day Cruise Intelligence Media Crawler
What the crawler is
Departure Day Cruise Intelligence operates a media monitoring system to understand how noteworthy events in the cruise industry propagate through the cruise media ecosystem.
We may link to external articles about cruise lines, ships, sailings, or cruise ports. We do not republish articles in whole or in part, and we do not use them to train AI models.
We may also use these media observations to notify users when breaking news is relevant to a sailing they are scheduled to travel on or a sailing they are watching.
Why we monitor public media
Departure Day uses public media observations to identify cruise industry posts and developments, group related posts concerning the same underlying event, and establish factual coverage timelines. This work can help show when our systems first observed a post, how coverage developed across publications, and which sources appear to be the earliest observed sources when the available evidence supports that conclusion.
We may also use the crawler to track and link to media posts about cruise lines, ships, or individual sailings, and to connect those posts with relevant cruise industry entities and events in Departure Day Cruise Intelligence.
What information may be observed
Depending on the source, observations may include a public page URL, page title, publication or source name, author or byline when available, publication and update timestamps presented by the source, the time Departure Day first or subsequently observed the page, headings, summaries, article text needed for analysis, links, and technical response information needed to verify the observation.
A “first observed” timestamp records when Departure Day’s systems first detected content. It is not necessarily the content’s original publication time. Departure Day may maintain historical observations needed to reconstruct how a post developed, including limited observations of public content that was later changed or removed. We do not publish a fixed retention period on this page.
How the information is used
Observed information may be used for search, classification, post clustering, source attribution, timeline reconstruction, historical comparison, verification, and analysis of how cruise industry coverage develops. It may also support links from Cruise Intelligence pages for lines, ships, or sailings to relevant reporting by source publications.
Post clusters and possible source relationships may be generated algorithmically and can be corrected when better evidence becomes available. A later article covering the same event does not, by itself, establish that one publication copied, relied upon, or learned of the event from another publication.
Crawler identification
The crawler identifies itself in HTTP requests with this User Agent:
Departure Day Cruise Intelligence Media Crawler/1.0 (+https://departureday.app/crawler)
Crawl behavior and request rates
The crawler is designed to make targeted requests to relevant public pages and avoid unreasonable traffic. Crawl frequency may vary by source and operational need. Departure Day does not publish a universal request frequency because sources differ and crawler scheduling may change.
The crawler does not attempt to reach authenticated, paywalled, or other content that is not public by circumventing access controls. Where a platform provides an official API suited to the task, Departure Day may use that API instead of directly crawling the platform’s pages.
robots.txt behavior
Before requesting a target page, the crawler fetches the applicable site’s
robots.txt file and evaluates the rules for its User Agent. It follows applicable
Allow and Disallow rules, along with publisher crawl delay instructions. Redirect
destinations are checked against the destination site’s robots rules before content is fetched.
If robots rules cannot be retrieved or evaluated, the crawler does not proceed with the target
request. A missing robots.txt file, represented by an HTTP 404 or 410 response, is
treated as having no published rules for that site.
Links and attribution
When Departure Day presents or references a media observation, it should identify the source publication and link to the originating public page when practical. Inclusion does not imply that a publisher has partnered with, endorsed, or authorized Departure Day.
Copyright and content handling
Departure Day’s system is intended to retain only the information and content necessary for discovery, classification, analysis, provenance, historical comparison, and verification of media timelines. The crawler is not intended to create a replacement reading experience for an originating publication.
Copyright remains with the applicable rights holders. Departure Day may use limited excerpts or internal observations where needed to identify, analyze, compare, and verify public coverage, and seeks to direct readers to the source publication for the original reporting.
Publisher contact
Publishers can contact Departure Day about crawling, attribution, technical issues, or information they believe is incorrect by emailing media@departureday.app. Please include the affected URL, the publication name, and enough detail for us to investigate.
Interpreting media observations
A first observed time shows when Departure Day first detected content. The source’s stated publication time, when available, is recorded separately.
Post clustering identifies related coverage. Any possible source relationship is an analytical assessment, not proof of copying or reliance, and may be corrected as better evidence becomes available.
Monitoring or linking to public coverage does not indicate a partnership, endorsement, or authorization by the publisher.