Sitemap: https://www.datanexus.ae/sitemap.xml # Data Nexus — https://www.datanexus.ae # Data Nexus Technologies Services FZ-LLC. Ras Al Khaimah Economic Zone Authority (RAKEZ) — # Media Licence No. 17009226 · Services Licence No. 47034164. # # Nothing on this host is blocked from any crawler. That is a decision rather # than a default, and it is worth writing down class by class. # # Search indexers — Googlebot, Bingbot, Applebot, YandexBot, DuckDuckBot. # They send readers, and they are also the assistant supply chain: a page # missing from Bing is missing from the answers that read Bing. Nothing # under /_next/ is blocked either — a robots.txt that hides the CSS and JS # from the renderer is a self-inflicted wound. # # Answer-engine indexers — OAI-SearchBot, Claude-SearchBot, PerplexityBot. # These build the retrieval index an assistant cites from. Blocking them is # the one unambiguous own goal available to a practice that sells search # and AI answer visibility. # # User-triggered fetchers — ChatGPT-User, Claude-User, Perplexity-User. # Not crawlers. Each fetches one URL because a person asked a question a # moment ago and is waiting. Blocking them means that person is told the # page could not be read. It is the highest-intent traffic this site will # ever receive. # # Training crawlers — GPTBot, ClaudeBot, CCBot, meta-externalagent, # cohere-ai, Amazonbot, Bytespider. # The genuinely contestable one, so the trade is stated rather than # inherited. Training returns no citation, no link and no visit. Allowed # anyway: this firm's product is the work, not the writing about it, and an # argument that is not in the weights is not in the answer given when # nothing is fetched at all. # # Training-use tokens — Google-Extended, Applebot-Extended. # These fetch nothing. They govern only whether pages already crawled may # be used for model training. Allowing them costs no crawl budget and has # no effect on ranking. Easy to leave blocked by accident. # # SEO and backlink crawlers — AhrefsBot, SemrushBot, DataForSeoBot. # Allowed. This practice tracks its own domain with these tools, and a # crawler you block is a report you cannot run. # # Archivers — archive.org_bot. # Allowed. A dated public record of what this site claimed is one more way # to check us, which is the argument the site makes about everything else. # # The Content-Signal line below says the same thing in the machine-readable form # the format defines: search=yes, ai-input=yes, ai-train=yes. Three yeses is an # unusual policy and it is deliberate. The format exists so a site can reserve # rights; a readiness checker marks the line present either way, which means the # score rewards declaring the field rather than declaring anything in # particular. Ours grants, because the prose above already does and a machine # should not have to read English to find that out. # # Content signals express a preference. They do not enforce one, and no operator # is obliged to honour them. # # Every agent named, and the rules written once. This file used to carry a # single wildcard group, on the argument that a named group REPLACES the # wildcard rather than adding to it — so naming thirty agents creates thirty # ways for a future Disallow to be skipped by exactly the crawlers that matter # most. The argument is correct and the conclusion was wrong twice over. # # It was wrong about the audience. A third-party audit read the single-group # form and reported Google-Extended as blocked on 121 of 122 pages. It is not: # the file has no Disallow, the server answers that agent with 200 and the full # page, no meta or header restricts it, and a standards parser agrees. But a # practice that sells answer-engine visibility cannot have a mainstream tool # telling its clients otherwise, and being right is not the same as being # checkable. # # It was wrong about the remedy. The hazard is that rules and groups drift # apart, and the fix for that is construction rather than restraint: the rules # below are one constant, the groups are generated from lib/crawlers.ts, and a # Disallow added tomorrow is written into every group by the same loop. The # Content-Signal line repeats for the same reason — a named group replaces the # wildcard, so an agent that has its own group would not otherwise see it. # # No Crawl-delay: forty static pages behind a CDN do not have a crawl-load # problem, and if one appears the control belongs at the CDN. # # No Disallow lines. There is an /api/ on this host now — the payment call, its # webhook, and the outbound-click beacon — and every one of them answers POST # only, so a crawler making the one request it knows how to make gets 405 and # nothing to index. A Disallow would not protect them; it would advertise three # endpoint addresses to the readers with the least business knowing them. The # pages not meant for the index carry a noindex in their own heads instead: # /waitlist, and the three payment outcomes under /pay/. # # For machine readers: # Curated map of this site: https://www.datanexus.ae/llms.txt # Sitemap: https://www.datanexus.ae/sitemap.xml # Feed, RSS 2.0: https://www.datanexus.ae/feed.xml # Feed, JSON Feed 1.1: https://www.datanexus.ae/feed.json # Questions about crawling, including rate: contact@datanexus.ae # # Generated by app/robots.txt/route.ts from lib/company.ts. Edits made to a # served copy are overwritten at build. User-agent: * Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: ChatGPT-User Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Claude-User Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Perplexity-User Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: MistralAI-User Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: DuckAssistBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Meta-ExternalFetcher Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: OAI-SearchBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Claude-SearchBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: PerplexityBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Applebot-Extended Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Applebot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Googlebot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Bingbot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: YandexBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: DuckDuckBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Baiduspider Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: GPTBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: ClaudeBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Google-Extended Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Google-CloudVertexBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Meta-ExternalAgent Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: CCBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Bytespider Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: Amazonbot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: cohere-ai Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: PetalBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: ia_archiver Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: archive.org_bot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: AhrefsBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: SemrushBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes User-agent: DataForSeoBot Allow: / Content-Signal: search=yes, ai-input=yes, ai-train=yes