Privacy & trust
What an AI Assistant May Do With This Site
The crawl rules have always been machine-readable. Now they are readable by a person too, with twenty agents named by hand and a clear answer to whether you may quote this.
Applies to Novus Learn 0.1.0
A policy file is not an answer
Copy linkA robots file has answered requests here for months. What did not exist was a page a person could read to find out whether they may quote this site, which agents are named, and what a crawl leaves behind. Telling someone it is in robots.txt is not an answer to a question about attribution, because a user-agent line says nothing about how the material may be used afterwards.
The AI access page is that answer, and it is generated from the same module that emits the robots rules. Retyping the agent list in prose is the obvious way to build such a page and the reason it would eventually lie: an agent added to the rules would be missing from the page claiming to describe them, and nothing would fail.
Two kinds of agent, treated differently
Copy linkThe distinction that matters is not which company operates an agent. It is whether a person is waiting. A bulk crawler reads broadly to build an index or a corpus, with nobody sitting in front of it. A user-initiated fetcher reads one page because somebody asked a question and is waiting for the answer, which makes it a reader being served.
Both kinds are named individually, operator by operator, and the page separates them into those two groups rather than listing every token in one block. It prints the current count from the same module the rules come from, so the figure on the page is the figure in force. The withheld paths differ between the groups for exactly this reason.
Why the tone here is permissive
Copy linkNovus Learn publishes educational material assembled from openly licensed public sources. Being readable by an assistant is close to the point of it: a learner who asks an assistant to explain something and is handed this site's explanation has been served, not robbed. So the policy asks for attribution rather than restricting access.
That is a decision specific to this site's content, not a general position. A site whose value is original reporting or a proprietary dataset has an entirely different calculation to make, and copying this policy across without that thought would be a mistake.
What is actually withheld, and why it is boring
Copy linkThe withheld list is short and contains nothing interesting, which is the point. There is no hidden section, no gated archive and no members area, because there are no members. What is kept out is kept out because serving it to a crawler would produce a page with no content in it.
One path differs between the two groups, and the page prints that difference rather than merging the rules. A bulk crawler is kept off the search tool, because crawling a search interface generates query pages rather than material; a fetcher acting for a person who asked a question is not, because that person may genuinely have meant to search.
- Device-local pages that read this browser's own storage and would return an empty shell to anyone else.
- Private result, attempt, session and share addresses, which are not linked from anywhere public.
- The search tool itself, for bulk crawlers only, because crawling a search interface produces query pages rather than content.
Two files that exist for the same reason
Copy linkAlongside the robots rules the site publishes two plain-text maps aimed at AI programs: a curated one naming the destinations that matter, and a complete enumeration of every indexable URL. Both are generated from the same registries as the tool map and the sitemap, so a program reading either of them sees the same site a person does.
That symmetry is the part worth insisting on. A site that serves one shape of itself to readers and another to machines has two things to keep true, and it will eventually fail to.
Attribution, in practical terms
Copy linkIf you quote a page here, name the page and link to it. If you reuse material that came from Wikimedia by way of this site, the original licence still governs it and the attribution requirements travel with the text, which is why every study export carries its source list rather than only its content.
The licences page separates three different questions that get merged too often: what you may do with the application, what you may do with the content this site authored, and what you may do with the brand. They are not under the same terms.
If you are writing about the product
Copy linkThe press kit holds the logo files, the boilerplate description and the facts that are safe to quote, together with what you may and may not do with them. Every asset listed there is a real file that ships with the site, which is checked rather than assumed.