{"author":"SpikeyCoder","children":[{"author":"ButlerianJihad","children":[{"author":"freedomben","children":[],"created_at":"2026-08-02T12:33:52.000Z","created_at_i":1785674032,"id":49143946,"options":[],"parent_id":49143918,"points":null,"story_id":49143630,"text":"For standard responses you&#x27;re not wrong, but increasingly there are a lot of LLMS that are essentially doing RAG against search results. For example this is I believe how Kagi works, and Google AI overviews. I have increasingly seen Claude and chat GPT also doing the same thing where instead of answering a question from the knowledge Bank they will do a web search and cite the responses that were used.","title":null,"type":"comment","url":null},{"author":"SecretDreams","children":[{"author":"pfdietz","children":[{"author":"SecretDreams","children":[],"created_at":"2026-08-02T12:50:56.000Z","created_at_i":1785675056,"id":49144129,"options":[],"parent_id":49144121,"points":null,"story_id":49143630,"text":"Yep..the online link posters hit close to home. Big reason I hardly even engage in discourse anymore. Got tired of verifying links that don&#x27;t have the claimed content in them.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T12:49:45.000Z","created_at_i":1785674985,"id":49144121,"options":[],"parent_id":49143970,"points":null,"story_id":49143630,"text":"It&#x27;s hilariously common online as well.  Bullshitter makes a claim, someone calls them on the BS, they put up a link, and (lo and behold) the link when examined does nothing to defend the claim.  Often the link directly contradicts the claim.","title":null,"type":"comment","url":null},{"author":"wwweston","children":[{"author":"SecretDreams","children":[],"created_at":"2026-08-02T15:47:25.000Z","created_at_i":1785685645,"id":49145630,"options":[],"parent_id":49145285,"points":null,"story_id":49143630,"text":"Yes, but it takes a lot of time to find a lie and very little time to conceive a lie. It&#x27;s a losing battle as the percentage of honest citations drops (and it doesn&#x27;t need to be a big drop to have a big impact.. think if the reliability of planes dropped 1%).","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T15:00:58.000Z","created_at_i":1785682858,"id":49145285,"options":[],"parent_id":49143970,"points":null,"story_id":49143630,"text":"with the advantage that you can discover this by reading citations.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T12:36:26.000Z","created_at_i":1785674186,"id":49143970,"options":[],"parent_id":49143918,"points":null,"story_id":49143630,"text":"Lol, this same behavior comes up more often than you&#x27;d expect in academia too. Many many citations that nobody cross checks that also don&#x27;t support the originally cited piece of information.","title":null,"type":"comment","url":null},{"author":"anon373839","children":[{"author":"captainbland","children":[{"author":"anon373839","children":[],"created_at":"2026-08-02T13:11:24.000Z","created_at_i":1785676284,"id":49144354,"options":[],"parent_id":49144291,"points":null,"story_id":49143630,"text":"Could they even afford 4B dense parameters touching every token? That seems like a LOT of compute compared to generating the classic Google SERP.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:06:41.000Z","created_at_i":1785676001,"id":49144291,"options":[],"parent_id":49144058,"points":null,"story_id":49143630,"text":"I&#x27;d be surprised if it were anything other than just an LLM only equivalent of Gemma some 4B param model, they can&#x27;t be spending much money on it because each query essentially needs to be cheaper than the ad revenue per search.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T12:44:41.000Z","created_at_i":1785674681,"id":49144058,"options":[],"parent_id":49143918,"points":null,"story_id":49143630,"text":"This is true, but good LLMs (even local, consumer-sized ones these days) can do this much better than \u2026 whatever it is that Google\u2019s AI overviews use.<p>Think about your coding harness: the model reads a bunch of files into the context, and it generally doesn\u2019t forget&#x2F;hallucinate which lines came from which file.","title":null,"type":"comment","url":null},{"author":"chollida1","children":[],"created_at":"2026-08-02T12:44:41.000Z","created_at_i":1785674681,"id":49144059,"options":[],"parent_id":49143918,"points":null,"story_id":49143630,"text":"&gt; It is architecturally impossible for an LLM to associate a link or citation that it crawled with a response that comes out the other end.<p>I mean, its not.  Lots of LLM&#x27;s focused on finance do this already.<p>Bloomberg&#x27;s own ASKB produces results and provides links back to the source documents or urls that it references so people who care about correctness can verify the results.","title":null,"type":"comment","url":null},{"author":"jefftk","children":[{"author":"knollimar","children":[],"created_at":"2026-08-02T13:13:44.000Z","created_at_i":1785676424,"id":49144383,"options":[],"parent_id":49144071,"points":null,"story_id":49143630,"text":"I don&#x27;t understand why verifying isn&#x27;t in the loop.  If you gave an LLM: &quot;is this statement consistent with this RAG data?&quot; it would do well.<p>It feels lazy that this isn&#x27;t built into the harness in some adversarial citation review checkbox.<p>Also the questions LLMs ask are so inhuman.  It&#x27;s like it&#x27;s some contrarian trying to win an argument over getting an answer.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T12:46:04.000Z","created_at_i":1785674764,"id":49144071,"options":[],"parent_id":49143918,"points":null,"story_id":49143630,"text":"You can do it the same way human does: recall a fact from memory, and then search to identify a citation. The citation isn&#x27;t &quot;here&#x27;s why I think this&quot; but instead &quot;here&#x27;s where you can verify this&quot;.","title":null,"type":"comment","url":null},{"author":"amelius","children":[{"author":"brookst","children":[{"author":"amelius","children":[],"created_at":"2026-08-02T13:20:08.000Z","created_at_i":1785676808,"id":49144444,"options":[],"parent_id":49144330,"points":null,"story_id":49143630,"text":"Yes. Whatever makes their lawyers nervous.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:09:32.000Z","created_at_i":1785676172,"id":49144330,"options":[],"parent_id":49144095,"points":null,"story_id":49143630,"text":"So an HTML page with a list of essentially every page on every site on the internet? What do you think, maybe a trillion URLs?","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T12:47:52.000Z","created_at_i":1785674872,"id":49144095,"options":[],"parent_id":49143918,"points":null,"story_id":49143630,"text":"Then the least they can do is cite ALL their sources. At least somewhere on their website.","title":null,"type":"comment","url":null},{"author":"WarmWash","children":[],"created_at":"2026-08-02T12:49:26.000Z","created_at_i":1785674966,"id":49144114,"options":[],"parent_id":49143918,"points":null,"story_id":49143630,"text":"This is why it&#x27;s dumb when inferencing with a model to ask &quot;...and cite all sources&quot;. Total waste of time and likely to make the response worse.<p>BUT<p>On models with web search, they can scan the sources in context, and those are pretty good at correct citations. However it&#x27;s still a &quot;trust, but verify&quot; situation.","title":null,"type":"comment","url":null},{"author":"ordersofmag","children":[],"created_at":"2026-08-02T13:00:03.000Z","created_at_i":1785675603,"id":49144218,"options":[],"parent_id":49143918,"points":null,"story_id":49143630,"text":"99% of AI users these days aren&#x27;t using a &#x27;raw&#x27; LLM.  They are interacting with a harness that includes tools to do search of the live (or recently crawled) web.  And so the output users actually see could absolutely include  correct citation of sources. So the &#x27;architecturally impossible&#x27; bit may be technically correct but it is not practically relevant.  Now whether those harnesses do a good job of orchestrating LLM output to get accurate citations (and whether they are transparent about their process) is another thing entirely.  But if you&#x27;re contemplating &#x27;what LLM&#x27;s can do&#x27; and aren&#x27;t taking into account the harness and tooling ecosystem they are embedded in then you&#x27;re missing the point.","title":null,"type":"comment","url":null},{"author":"ssl-3","children":[],"created_at":"2026-08-02T13:02:05.000Z","created_at_i":1785675725,"id":49144242,"options":[],"parent_id":49143918,"points":null,"story_id":49143630,"text":"&gt; It is perfectly common to find that a &quot;source&quot; does not contain the statements.<p>It is commonly this way right now, but it doesn&#x27;t have to stay this way forever.<p>It is also common, in my usage at least, to iteratively brow-beat the bot into paring its statements down to those that which are supportable by its sources.  Doing so just takes repetition, and that repetition takes time and burns more tokens.<p>With the present state of things, the prompts to get moving on this and to guide the ultimate response into something that is verifiably supportable by outside sources can usually be simple and largely generic.<p>They&#x27;re easy enough prompts that a subagent can produce them.<p>(I&#x27;ve done it myself with Codex subagents and it worked very well, aside from the unsustainable burn rate that did not fit my budget.)","title":null,"type":"comment","url":null},{"author":"razodactyl","children":[],"created_at":"2026-08-02T13:15:55.000Z","created_at_i":1785676555,"id":49144401,"options":[],"parent_id":49143918,"points":null,"story_id":49143630,"text":"Are you sure? Because this sounds like GPT3-era understandings of LLM-isms","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T12:30:34.000Z","created_at_i":1785673834,"id":49143918,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"It is architecturally impossible for an LLM to associate a link or citation that it crawled with a response that comes out the other end. Every link they&#x27;re giving you to support their statements is tacked on because it may vaguely match the tokens it just generated. It is perfectly common to find that a &quot;source&quot; does not contain the statements. You cannot expect an LLM to write you a Wikipedia article, much less a legal or medical opinion supported by research.","title":null,"type":"comment","url":null},{"author":"eterm","children":[{"author":"SpikeyCoder","children":[],"created_at":"2026-08-02T12:41:12.000Z","created_at_i":1785674472,"id":49144017,"options":[],"parent_id":49143953,"points":null,"story_id":49143630,"text":"I get that pushback. My goal with this index&#x2F;report wasn&#x27;t to write a new &#x27;AI SEO&#x27; playbook or encourage people to start gaming the system. The goal was simply to shine a bit of light on what&#x27;s going on. Right now, there is a massive information asymmetry: AI companies are turning their assistants into primary search engines, but webmasters have zero visibility into whether their sites are actually being cited in those live answers.<p>You&#x27;re absolutely right to be concerned about an arms race. If the data showed that doing X, Y, and Z guaranteed a citation, we&#x27;d be right back to the worst days of keyword stuffing.<p>But what this data actually shows is the opposite: even if you do everything &#x27;right&#x27; (allow the retrieval bots, provide perfect machine-readable schema, don&#x27;t block anything), you still have a 94.8% chance of never being cited. How LLMs cite is entirely opaque.<p>I built this to give site owners a baseline measurement of what is actually happening to their content today (as a free, anonymous view), so they can make an informed decision on whether keeping their doors open to these crawlers is actually worth it.","title":null,"type":"comment","url":null},{"author":"passwordoops","children":[],"created_at":"2026-08-02T12:57:50.000Z","created_at_i":1785675470,"id":49144188,"options":[],"parent_id":49143953,"points":null,"story_id":49143630,"text":"LLM&#x27;s were obviously going be ruined by two factors:<p>1- LLM Crawler Optimization: the new SEO<p>2- Weightings-For-Pay: for a fee have your product or service come up more frequently in associated answers","title":null,"type":"comment","url":null},{"author":"PunchyHamster","children":[],"created_at":"2026-08-02T13:23:21.000Z","created_at_i":1785677001,"id":49144467,"options":[],"parent_id":49143953,"points":null,"story_id":49143630,"text":"&gt; Are we saying that it&#x27;s now a problem that we&#x27;re not getting scraped?<p>No, it is saying that even the ones that are being scraped are not being included in results","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T12:34:17.000Z","created_at_i":1785674057,"id":49143953,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"Are we saying that it&#x27;s now a problem that we&#x27;re <i>not</i> getting scraped?<p>This appears to be a new generation of &quot;SEO&quot;, marking itself as a service for getting into AI results?<p>This is not the future I want to be a part of.<p>Perhaps it&#x27;s inevitable that after a break from everything being driven by money that LLMs will now be ruined by people spending $X to get into AI to make back $X+1, leading to an arms race of ever increasing X, to the detriment of users.","title":null,"type":"comment","url":null},{"author":"josh-wrale","children":[],"created_at":"2026-08-02T12:43:09.000Z","created_at_i":1785674589,"id":49144048,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"Makes me wonder if paid Medium and paid news are input to AI training. Surely, they are.","title":null,"type":"comment","url":null},{"author":"pfdietz","children":[],"created_at":"2026-08-02T12:48:16.000Z","created_at_i":1785674896,"id":49144100,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"&gt; 94.8%<p>This seems consistent with Sturgeon&#x27;s Law.","title":null,"type":"comment","url":null},{"author":"mohamedkoubaa","children":[],"created_at":"2026-08-02T12:48:42.000Z","created_at_i":1785674922,"id":49144104,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"We don&#x27;t have a page rank mechanism. Most likely people are paying or threatening AI companies to boost certain sources as authoritative","title":null,"type":"comment","url":null},{"author":"lorreyfum","children":[],"created_at":"2026-08-02T12:53:19.000Z","created_at_i":1785675199,"id":49144148,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"Call it web 4.0","title":null,"type":"comment","url":null},{"author":"brookst","children":[{"author":"DrDeese","children":[],"created_at":"2026-08-02T12:56:44.000Z","created_at_i":1785675404,"id":49144176,"options":[],"parent_id":49144165,"points":null,"story_id":49143630,"text":"I wouldn&#x27;t even consider it a complaint. For a builder, I would consider it an opportunity...","title":null,"type":"comment","url":null},{"author":"kyleblarson","children":[{"author":"embedding-shape","children":[],"created_at":"2026-08-02T13:09:27.000Z","created_at_i":1785676167,"id":49144327,"options":[],"parent_id":49144238,"points":null,"story_id":49143630,"text":"That&#x27;s a &quot;problem&quot; of the harness that you use those, not a question of the model. The model can &quot;know&quot; things like &quot;Linus Torvalds created Linux&quot; but things like &quot;Is X available in Y right now?&quot; it obviously cannot, leading to this being a completely different thing compared to &quot;just knowing&quot; something and providing references to &quot;how it knows&quot;.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:01:28.000Z","created_at_i":1785675688,"id":49144238,"options":[],"parent_id":49144165,"points":null,"story_id":49143630,"text":"The question is very important. What if you ask&#x2F;tell an AI &quot;I need a 2 bedroom vacation rental in Park City for a trip this fall&quot;. If you only get Airbnb results you are missing a lot of the true answer.","title":null,"type":"comment","url":null},{"author":"SpikeyCoder","children":[{"author":"johndhi","children":[{"author":"SpikeyCoder","children":[],"created_at":"2026-08-02T13:12:14.000Z","created_at_i":1785676334,"id":49144368,"options":[],"parent_id":49144290,"points":null,"story_id":49143630,"text":"Small businesses don&#x27;t. You&#x27;re right. But somehow those larger companies seem to get cited. Wouldn&#x27;t it be nice if we all knew how we were being influenced, both potential consumer and small business owner? Let&#x27;s have the lid off!","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:06:33.000Z","created_at_i":1785675993,"id":49144290,"options":[],"parent_id":49144247,"points":null,"story_id":49143630,"text":"Makes sense! Good point.<p>You&#x27;re effectively reminding us that brands don&#x27;t get understand well how to influence ai seo","title":null,"type":"comment","url":null},{"author":"spiderfarmer","children":[],"created_at":"2026-08-02T13:07:05.000Z","created_at_i":1785676025,"id":49144299,"options":[],"parent_id":49144247,"points":null,"story_id":49143630,"text":"That\u2019s why I, from now on, only allow citable information to be scraped. The return of cloaking, lol.","title":null,"type":"comment","url":null},{"author":"altmanaltman","children":[{"author":"martinald","children":[],"created_at":"2026-08-02T13:25:26.000Z","created_at_i":1785677126,"id":49144489,"options":[],"parent_id":49144384,"points":null,"story_id":49143630,"text":"It depends. If you are doing blog content with the idea of upselling users to your product, probably not as much (because the LLM can just give the user the answer).<p>However, if you are looking for the best product&#x2F;service&#x2F;whatever, then yes it really does matter. I&#x27;ve bought _so_ many products because of LLM recommendations. For example, I wanted a new webcam, I asked the LLM to find me the best ones with a large sensor and Linux compatibility. It gave me a shortlist then I chose one and then I bought it.<p>This experience is far better than trailing through dozens of pages of (even pre LLM) SEO slop.I just tried the same on Google search and all the links recommended a camera with ~10% the sensor size that I bought.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:13:47.000Z","created_at_i":1785676427,"id":49144384,"options":[],"parent_id":49144247,"points":null,"story_id":49143630,"text":"Not really, ahrefs has already started working on this and there are several tools out there that track GEO.<p>Also what is the incentive for me to rank on LLMs? People do not use it to click sources or even take action. Do you have any proof that being ranked as the best roofer in houston drives sales more than not being ranked as it? If not, why should i care as the roofer?","title":null,"type":"comment","url":null},{"author":"martinald","children":[],"created_at":"2026-08-02T13:19:33.000Z","created_at_i":1785676773,"id":49144438,"options":[],"parent_id":49144247,"points":null,"story_id":49143630,"text":"That&#x27;s not _entirely_ true. Search Console now has a Generative AI page where you can see impressions per page now. Bing has something similar.<p>Also, I would assume that really &quot;GEO&quot; is just like &quot;SEO&quot;. If you rank on the &#x27;1st page&#x27; of results for whatever common searches, you are _very_ likely to rank the same way for LLM questions, because all the LLM is doing (nearly all of the time for &#x27;best roofers in Houston&#x27;) is doing a web search and summarising the first x results. So if you are on page 1 for that term, it&#x27;s very likely IME that you will get mentioned on LLM answers for that.","title":null,"type":"comment","url":null},{"author":"phoghed","children":[{"author":"philipallstar","children":[{"author":"orbital-decay","children":[],"created_at":"2026-08-02T13:41:51.000Z","created_at_i":1785678111,"id":49144632,"options":[],"parent_id":49144545,"points":null,"story_id":49143630,"text":"The model does it in this case, not &quot;we&quot;. Classic SEO does influence it because it relies on traditional search, but then it estimates the best according to whatever cognitive capabilities it has, to its understanding of user&#x27;s request, to what its harness tells it to do, and to its character engineered during post-training (and models&#x27; biases are <i>very carefully</i> engineered).<p>I would expect a judgement of a modern model to be superficial for this request, but still better than nothing. A deep research agent might be a lot better and much less susceptible to &quot;optimization&quot;.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:32:15.000Z","created_at_i":1785677535,"id":49144545,"options":[],"parent_id":49144451,"points":null,"story_id":49143630,"text":"How are we defining the best roofers?","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:21:09.000Z","created_at_i":1785676869,"id":49144451,"options":[],"parent_id":49144247,"points":null,"story_id":49143630,"text":"Good, let\u2019s hope there\u2019s no SEO arms race for this too. Maybe the best roofers have a chance of getting cited over the ones with the best SEO agency. Seems unlikely though.","title":null,"type":"comment","url":null},{"author":"orbital-decay","children":[],"created_at":"2026-08-02T13:25:01.000Z","created_at_i":1785677101,"id":49144480,"options":[],"parent_id":49144247,"points":null,"story_id":49143630,"text":"<i>&gt;But if I ask an LLM &#x27;who are the best commercial roofers in Houston?&#x27;, there isn&#x27;t one right answer. When an LLM answers that roofing question, it typically cites 3 to 5 businesses. The other 40 legitimate roofing companies in the area are left out.</i><p>I don&#x27;t see how mentioning the other 40 should follow from that fact. You get exactly what you&#x27;re asking for - the best, according to agent&#x27;s judgement. Judgement and processing of simple search results is exactly what you&#x27;re using the intelligent agent for. Ask for a <i>list</i> of all commercial roofers in Houston and then you can expect a list.<p>It can be argued that models are clustering their replies around a few options when asked to choose from a list of equally valid ones (due to mode collapse, distribution biases, primacy&#x2F;recency biases etc), but these options are not equally valid for the agent, it looks at their sites or possibly in some other places like review sites to rank them. Its ranking criteria might be not good at all, but that&#x27;s another question.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:02:22.000Z","created_at_i":1785675742,"id":49144247,"options":[],"parent_id":49144165,"points":null,"story_id":49143630,"text":"If the goal of the web was just to transmit objective facts like &#x27;who created Linux&#x27;, this wouldn&#x27;t be a problem at all. Wikipedia handles that pretty dece.<p>The issue arises when we move away from objective trivia and into subjective, localized buying intent, which is where the web monetizes itself.<p>If I ask an LLM &#x27;who created Linux?&#x27;, there is one right answer. But if I ask an LLM &#x27;who are the best commercial roofers in Houston?&#x27;, there isn&#x27;t one right answer. There are dozens of highly qualified local businesses that do possess unique value, unique pricing, and unique availability.<p>When an LLM answers that roofing question, it typically cites 3 to 5 businesses. The other 40 legitimate roofing companies in the area are left out. My study isn&#x27;t arguing that every single one of those 40 companies deserves to be in the answer; it&#x27;s pointing out that those 40 companies currently have no idea they are being left out.<p>In the Google era, if you weren&#x27;t on page 1, you could look at Search Console, see your ranking, check your backlinks, and understand why. In the AI Search era, businesses are being scraped to build these answers, but they have zero telemetry on whether they are actually making the cut. This index is just an attempt to provide that missing telemetry.","title":null,"type":"comment","url":null},{"author":"tiffanyh","children":[],"created_at":"2026-08-02T13:03:57.000Z","created_at_i":1785675837,"id":49144266,"options":[],"parent_id":49144165,"points":null,"story_id":49143630,"text":"I just asked ChatGPT-5.6 that question, no source was given.<p>Not even Wikipedia.","title":null,"type":"comment","url":null},{"author":"altmanaltman","children":[{"author":"SpikeyCoder","children":[],"created_at":"2026-08-02T13:16:22.000Z","created_at_i":1785676582,"id":49144408,"options":[],"parent_id":49144351,"points":null,"story_id":49143630,"text":"I disclosed on my first comment how I got my data. I thought it was an interesting find, and one that really hinders the non-technical small business owner from figuring out how to ensure their company finds their way to the right customer.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:11:19.000Z","created_at_i":1785676279,"id":49144351,"options":[],"parent_id":49144165,"points":null,"story_id":49143630,"text":"Its a marketing post for a tool that offers visibility to site owners. Honestly these type of &quot;neutral data reports&quot; masquerades should be flagged or adequately disclosed.","title":null,"type":"comment","url":null},{"author":"llm_nerd","children":[],"created_at":"2026-08-02T13:15:38.000Z","created_at_i":1785676538,"id":49144399,"options":[],"parent_id":49144165,"points":null,"story_id":49143630,"text":"It isn&#x27;t a complaint. It&#x27;s an ad. This is literally an ad for some sort of &quot;get cited by AI&quot; service.","title":null,"type":"comment","url":null},{"author":"netcan","children":[],"created_at":"2026-08-02T13:23:17.000Z","created_at_i":1785676997,"id":49144466,"options":[],"parent_id":49144165,"points":null,"story_id":49143630,"text":"This is not a complaint.<p>I think your assessment of this result being as expected... but this is about the LLM equivalent of SEO becoming an area of interest","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T12:55:23.000Z","created_at_i":1785675323,"id":49144165,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"This is a weird complaint.<p>Let\u2019s say I ask \u201cwho created Linux?\u201d Claude correctly tells me Linus Torvalds, and links to Wikipedia.<p>There are probably thousands, maybe hundreds of thousands of other sites that have that some piece of information. Are LLMs supposed to link to <i>every</i> site?<p>Most sites do not have unique information at all, and even sites that do rarely contain <i>only</i> unique information. Implying that most sites deserve links because they were crawled seems like a statistical fallacy. You could say the same thing about the percent of sites crawled by Google versus ever showing up on first page of results.","title":null,"type":"comment","url":null},{"author":"sparkling","children":[],"created_at":"2026-08-02T12:59:03.000Z","created_at_i":1785675543,"id":49144206,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"Blocking known bot identifiers via robots.txt does nothing by the way. Too many labs are running sneaky crawlers that do not respect robots.txt. You will need to take extreme measures: blocking basically all datacenter IP ranges, VPN IPs, aggressive rate limiting, etc.<p>Blocking LLM crawlers has become the number one use case for our IP database  customers at <a href=\"https:&#x2F;&#x2F;focsec.com&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;focsec.com&#x2F;</a>","title":null,"type":"comment","url":null},{"author":"ermantrout","children":[],"created_at":"2026-08-02T13:00:45.000Z","created_at_i":1785675645,"id":49144225,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"This makes sense. Why would everyone be coted for the same answer","title":null,"type":"comment","url":null},{"author":"mark_l_watson","children":[],"created_at":"2026-08-02T13:01:18.000Z","created_at_i":1785675678,"id":49144236,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"Odd results for me. Last month I tested five AI chat sites, with web search turned off, and four of them had a shadow of information about me as a person (what kind of books I write, what tech I use, and a random bit of other information). The linked site gave me a zero score because it was testing if the AI models recommended my site for business or sales queries.","title":null,"type":"comment","url":null},{"author":"alsetmusic","children":[],"created_at":"2026-08-02T13:09:16.000Z","created_at_i":1785676156,"id":49144324,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"This makes sense. A handful of websites hold most of the \u201ctrusted\u201d info because they\u2019re massive. I don\u2019t expect you to quote my blog with only three entries. The real trick would be getting AI companies to stop hammering sites that don\u2019t show up in answers, but even if they don\u2019t use a source for an answer, crawling still provides value.","title":null,"type":"comment","url":null},{"author":"deadbabe","children":[{"author":"phoghed","children":[{"author":"deadbabe","children":[],"created_at":"2026-08-02T16:12:35.000Z","created_at_i":1785687155,"id":49145836,"options":[],"parent_id":49144486,"points":null,"story_id":49143630,"text":"Do you want to be cited, or do you want to make money?","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:25:22.000Z","created_at_i":1785677122,"id":49144486,"options":[],"parent_id":49144403,"points":null,"story_id":49143630,"text":"Seems like the opposite if the problem you\u2019re trying to solve is \u201cmy site isn\u2019t cited over someone else\u2019s\u201d. HTTP 467 - pls cite me, I\u2019ll pay you.","title":null,"type":"comment","url":null},{"author":"ndriscoll","children":[{"author":"deadbabe","children":[{"author":"ndriscoll","children":[],"created_at":"2026-08-02T16:59:02.000Z","created_at_i":1785689942,"id":49146198,"options":[],"parent_id":49145849,"points":null,"story_id":49143630,"text":"In a sense, that already exists. e.g. spotify or youtube premium. In practice only extremely exceptional individuals are going to see anything from it because creative markets are oversaturated, it&#x27;s trivial to enter them, and people have finite time, so almost everyone is just not worth anyone&#x27;s attention, let alone compensation.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T16:14:15.000Z","created_at_i":1785687255,"id":49145849,"options":[],"parent_id":49144624,"points":null,"story_id":49143630,"text":"Yes but imagine a platform instead where watching any YouTube video is paid for by a micro transaction. Transparently.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:40:01.000Z","created_at_i":1785678001,"id":49144624,"options":[],"parent_id":49144403,"points":null,"story_id":49143630,"text":"The problems with micropayments are<p>1. The market for lemons. In fact sites trying to &quot;monetize their content&quot; are the most likely to be lemons, so just asking is a signal that your &quot;content&quot; is not worth anything.<p>2. If your information does have value, it competes in a market with other sites full of high quality information that aren&#x27;t trying to monetize it, meaning your specific information has to be specifically very valuable and not available yet for free elsewhere. If it is valuable as information (i.e. not something like creative writing), it will quickly spread and become freely available.<p>Lots of &quot;content&quot; just isn&#x27;t worth anything (or has negative worth: it wastes your time). e.g. consider youtubers begging to get viewers to like&#x2F;subscribe to increase their reach, and people generally don&#x27;t despite it costing them nothing. Because it&#x27;s not even worth a click to them.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T13:16:01.000Z","created_at_i":1785676561,"id":49144403,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"This is why we need the HTTP 402 standard to become common.<p>If websites charge pennies per AI crawl, they will make more money than ever being reference in that 6.2% of websites that get cited (of which even another small percent get any follow through that leads to a sale or ad click)<p>HTTP 402 also basically extends the pay per token model people have gotten used to with AI model providers, except applied to the whole web, with the added privacy benefit in that there is no need for sellers of content to \u201cknow their customer\u201d, and indeed it may  even be impossible to do so because of how the payment gateways operate.","title":null,"type":"comment","url":null},{"author":"Rabbit504030201","children":[],"created_at":"2026-08-02T13:34:23.000Z","created_at_i":1785677663,"id":49144565,"options":[],"parent_id":49143630,"points":null,"story_id":49143630,"text":"Nice study :) Keep updating!<p>The title is very misleading - authors openly state the low and slewed sample pool.","title":null,"type":"comment","url":null}],"created_at":"2026-08-02T11:53:52.000Z","created_at_i":1785671632,"id":49143630,"options":[],"parent_id":null,"points":40,"story_id":49143630,"text":null,"title":"Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers","type":"story","url":"https://website-auditor.io/ai-visibility-index"}
