{"author":"raunakchowdhuri","children":[{"author":"skadamat","children":[{"author":"adit_a","children":[],"created_at":"2026-04-06T17:59:49.000Z","created_at_i":1775498389,"id":47664477,"options":[],"parent_id":47663942,"points":null,"story_id":47662833,"text":"We&#x27;re releasing an open dataset for challenging structured extraction tasks as a starting point for people to do any comparisons soon!<p>vikp and the Datalab team have done great work in the space, but their structured extraction product is closer to our baseline &#x2F;extract api since both of those are single pass extractions.<p>Deep Extract is more accurate than any structured extraction product we&#x27;ve tried, <i>but</i> the approach comes with a very clear cost&#x2F;latency tradeoff over a single pass extraction. We have free credits if you&#x27;d like to do a side by side","title":null,"type":"comment","url":null}],"created_at":"2026-04-06T17:21:54.000Z","created_at_i":1775496114,"id":47663942,"options":[],"parent_id":47662833,"points":null,"story_id":47662833,"text":"How does this compare to DataLab (<a href=\"https:&#x2F;&#x2F;www.datalab.to&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;www.datalab.to&#x2F;</a>)","title":null,"type":"comment","url":null},{"author":"willwjack","children":[{"author":"raunakchowdhuri","children":[],"created_at":"2026-04-07T00:04:29.000Z","created_at_i":1775520269,"id":47669037,"options":[],"parent_id":47664615,"points":null,"story_id":47662833,"text":"The big one is that LLMs get lazy on repetitive tasks. They&#x27;ll skip rows or consolidate entries instead of grinding through every last one. So you need verify-and-re-extract loops rather than single-pass processing. Breaking work into sub-agent chunks with explicit correctness criteria defined upfront (e.g., &quot;line items must sum to the stated total&quot;) lets the system self-verify autonomously. At scale (28M+ fields), this approach actually outperformed expert human labelers!","title":null,"type":"comment","url":null}],"created_at":"2026-04-06T18:08:47.000Z","created_at_i":1775498927,"id":47664615,"options":[],"parent_id":47662833,"points":null,"story_id":47662833,"text":"Any learnings from deploying agents at such massive scale?","title":null,"type":"comment","url":null},{"author":"aleks5678","children":[{"author":"raunakchowdhuri","children":[],"created_at":"2026-04-06T23:55:42.000Z","created_at_i":1775519742,"id":47668975,"options":[],"parent_id":47666492,"points":null,"story_id":47662833,"text":"We&#x27;ve made a lot of changes in the past few months that make our standard extract much, much better, as well as Deep Extract for documents even longer than that. We&#x27;d love for you to give it a try!","title":null,"type":"comment","url":null}],"created_at":"2026-04-06T20:21:15.000Z","created_at_i":1775506875,"id":47666492,"options":[],"parent_id":47662833,"points":null,"story_id":47662833,"text":"We used Reducto and it did struggle with long documents.  As we process financial documents going over 300+ pages  using Gemini 3 Flash is producing high accuracy extracts super fast.","title":null,"type":"comment","url":null},{"author":"cyanydeez","children":[{"author":"observationist","children":[],"created_at":"2026-04-06T22:52:30.000Z","created_at_i":1775515950,"id":47668420,"options":[],"parent_id":47667171,"points":null,"story_id":47662833,"text":"It&#x27;s a good harness, sir.","title":null,"type":"comment","url":null}],"created_at":"2026-04-06T21:10:53.000Z","created_at_i":1775509853,"id":47667171,"options":[],"parent_id":47662833,"points":null,"story_id":47662833,"text":"I like to play guess which LLM open source package is that XKCD comic.<p>Looks like it&#x27;s something like: <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;docs&#x2F;transformers&#x2F;model_doc&#x2F;layoutxlm\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;docs&#x2F;transformers&#x2F;model_doc&#x2F;layoutxlm</a>","title":null,"type":"comment","url":null},{"author":"nbnn","children":[],"created_at":"2026-04-07T12:44:42.000Z","created_at_i":1775565882,"id":47674438,"options":[],"parent_id":47662833,"points":null,"story_id":47662833,"text":"Irud","title":null,"type":"comment","url":null}],"created_at":"2026-04-06T16:13:47.000Z","created_at_i":1775492027,"id":47662833,"options":[],"parent_id":null,"points":50,"story_id":47662833,"text":null,"title":"Reducto releases Deep Extract","type":"story","url":"https://reducto.ai/blog/reducto-deep-extract-agent"}
