{"author":"thm","children":[{"author":"esafak","children":[{"author":"dang","children":[],"created_at":"2024-04-01T19:23:57.000Z","created_at_i":1711999437,"id":39898051,"options":[],"parent_id":39897959,"points":null,"story_id":39896923,"text":"Ok we&#x27;ve taken deep document understanding out of the title above. Thanks!","title":null,"type":"comment","url":null},{"author":"kergonath","children":[],"created_at":"2024-04-01T22:42:52.000Z","created_at_i":1712011372,"id":39900261,"options":[],"parent_id":39897959,"points":null,"story_id":39896923,"text":"I am curious about the performance of their OCR and layout and table detection. Hopefully it\u2019s on par with Amazon, Google, or Microsoft\u2019s tools.","title":null,"type":"comment","url":null},{"author":"snats","children":[],"created_at":"2024-04-01T23:11:33.000Z","created_at_i":1712013093,"id":39900480,"options":[],"parent_id":39897959,"points":null,"story_id":39896923,"text":"A bit off topic but every time I see DTrOCR I remember that marketing is a good idea because DTrOCR[1] and Fuyu[2] are basically the same architecture[3].<p>[1] <a href=\"https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2308.15996v1.pdf\" rel=\"nofollow\">https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;2308.15996v1.pdf</a>\n[2] <a href=\"https:&#x2F;&#x2F;www.adept.ai&#x2F;blog&#x2F;fuyu-8b\" rel=\"nofollow\">https:&#x2F;&#x2F;www.adept.ai&#x2F;blog&#x2F;fuyu-8b</a>\n[3] If you don&#x27;t want to search for the figures I made a tiny post about it on my weblog: <a href=\"https:&#x2F;&#x2F;weblog.snats.xyz&#x2F;posts&#x2F;2024&#x2F;02&#x2F;16&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;weblog.snats.xyz&#x2F;posts&#x2F;2024&#x2F;02&#x2F;16&#x2F;</a>","title":null,"type":"comment","url":null},{"author":"vissidarte_choi","children":[],"created_at":"2024-04-02T04:45:24.000Z","created_at_i":1712033124,"id":39902470,"options":[],"parent_id":39897959,"points":null,"story_id":39896923,"text":"RAGFlow uses Yolov8 for its OCR&#x2F;layout recognition&#x2F;TSR(table structure recognition).  And RAGFlow uses large amount private data to train these models for them to perform well in some specialized scenarios.","title":null,"type":"comment","url":null},{"author":"yingfeng","children":[],"created_at":"2024-04-02T05:25:58.000Z","created_at_i":1712035558,"id":39902662,"options":[],"parent_id":39897959,"points":null,"story_id":39896923,"text":"We&#x27;ve used YOLOv8 as the object detection model, and use some public datasets, such as PubTable, CDLA, together with some private data to train the model. The model on Huggingface is the one trained using public dataset, and we would open this work later. We use YOLOv8 just because we want to let the document parser run without GPU, I think you could also try any other object detection models such as Detectron, and use the public datasets to train the model as well. We&#x27;ve not used transfomers for this task, because given limited data, it could not outperform traditional CNN based models.","title":null,"type":"comment","url":null}],"created_at":"2024-04-01T19:14:19.000Z","created_at_i":1711998859,"id":39897959,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"Apparently &quot;deep document understanding&quot; refers to OCR and structured document parsing: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;blob&#x2F;main&#x2F;deepdoc&#x2F;README.md\">https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;blob&#x2F;main&#x2F;deepdoc&#x2F;READ...</a><p>Since &quot;deep document understanding&quot; is not a term of art, I would have just said &quot;OCR and document parsing&quot;.<p>How well does it work? Please include benchmarks. You may be interested in<p><a href=\"https:&#x2F;&#x2F;paperswithcode.com&#x2F;sota&#x2F;optical-character-recognition-on-benchmarking\" rel=\"nofollow\">https:&#x2F;&#x2F;paperswithcode.com&#x2F;sota&#x2F;optical-character-recognitio...</a><p><a href=\"https:&#x2F;&#x2F;paperswithcode.com&#x2F;task&#x2F;document-layout-analysis\" rel=\"nofollow\">https:&#x2F;&#x2F;paperswithcode.com&#x2F;task&#x2F;document-layout-analysis</a><p>The models seem to be closed source, hosted here: <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;InfiniFlow&#x2F;deepdoc\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;InfiniFlow&#x2F;deepdoc</a>","title":null,"type":"comment","url":null},{"author":"gardenfelder","children":[{"author":"shekhar101","children":[{"author":"gardenfelder","children":[],"created_at":"2024-04-02T14:42:12.000Z","created_at_i":1712068932,"id":39906277,"options":[],"parent_id":39898516,"points":null,"story_id":39896923,"text":"<a href=\"https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;liteLLM-proxy\">https:&#x2F;&#x2F;github.com&#x2F;BerriAI&#x2F;liteLLM-proxy</a>","title":null,"type":"comment","url":null}],"created_at":"2024-04-01T20:02:19.000Z","created_at_i":1712001739,"id":39898516,"options":[],"parent_id":39898012,"points":null,"story_id":39896923,"text":"It&#x27;s trivial to run a proxy server that routes all OpenAi calls to another LLM, even local ones. See litellm-proxy.","title":null,"type":"comment","url":null},{"author":"_akhe","children":[{"author":"gardenfelder","children":[{"author":"_akhe","children":[],"created_at":"2024-04-03T19:47:59.000Z","created_at_i":1712173679,"id":39922164,"options":[],"parent_id":39906414,"points":null,"story_id":39896923,"text":"<a href=\"https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;blob&#x2F;main&#x2F;rag&#x2F;llm&#x2F;chat_model.py#L132\">https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;blob&#x2F;main&#x2F;rag&#x2F;llm&#x2F;chat...</a>","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T14:53:57.000Z","created_at_i":1712069637,"id":39906414,"options":[],"parent_id":39899361,"points":null,"story_id":39896923,"text":"where do you see that?","title":null,"type":"comment","url":null}],"created_at":"2024-04-01T21:08:36.000Z","created_at_i":1712005716,"id":39899361,"options":[],"parent_id":39898012,"points":null,"story_id":39896923,"text":"I see a `LocalLLM` chat model where it looks like you can pass a host&#x2F;port (for example, ollama&#x27;s)","title":null,"type":"comment","url":null},{"author":"rosspackard","children":[],"created_at":"2024-04-01T23:08:31.000Z","created_at_i":1712012911,"id":39900457,"options":[],"parent_id":39898012,"points":null,"story_id":39896923,"text":"Looks like they do but aren&#x27;t really documented yet:<p><a href=\"https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;pull&#x2F;119\">https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;pull&#x2F;119</a>","title":null,"type":"comment","url":null},{"author":"vissidarte_choi","children":[],"created_at":"2024-04-02T04:48:08.000Z","created_at_i":1712033288,"id":39902485,"options":[],"parent_id":39898012,"points":null,"story_id":39896923,"text":"RAGFlow will support more LLMs, including locally deployed LLMs.","title":null,"type":"comment","url":null}],"created_at":"2024-04-01T19:20:12.000Z","created_at_i":1711999212,"id":39898012,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"It seems to be limited to certain LLM servers, on of which is OpenAI, none of which includes e.g. Mystral and popular OSS LLMs.<p>I wonder if that will change - eventually.<p>Discord channels are named in Chinese, though there are English posts.","title":null,"type":"comment","url":null},{"author":"_akhe","children":[{"author":"vissidarte_choi","children":[],"created_at":"2024-04-02T02:45:43.000Z","created_at_i":1712025943,"id":39901893,"options":[],"parent_id":39899351,"points":null,"story_id":39896923,"text":"Hi bschmidt1, This is a good feature. We do plan to support it soon. Please stay tuned. If you have further suggesions, welcome to file an issue with us.","title":null,"type":"comment","url":null}],"created_at":"2024-04-01T21:07:48.000Z","created_at_i":1712005668,"id":39899351,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"Is there a JavaScript library? Both LlamaIndex and Langchain have nice JS&#x2F;TS packages on npm. Could thinly wrap a JS client around this Python API but the community aspect of having an official library is nice.<p>Also might be helpful to have a simple example on the README showing how to fetch a document and start querying it. I would try it!","title":null,"type":"comment","url":null},{"author":"NKosmatos","children":[{"author":"rosspackard","children":[],"created_at":"2024-04-01T23:08:14.000Z","created_at_i":1712012894,"id":39900451,"options":[],"parent_id":39900324,"points":null,"story_id":39896923,"text":"Looks like they do but aren&#x27;t really documented yet:<p><a href=\"https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;pull&#x2F;119\">https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;pull&#x2F;119</a>","title":null,"type":"comment","url":null},{"author":"vissidarte_choi","children":[],"created_at":"2024-04-02T04:51:53.000Z","created_at_i":1712033513,"id":39902496,"options":[],"parent_id":39900324,"points":null,"story_id":39896923,"text":"To be honest, RAGFlow already supports this but has not documented this local deployment process yet, as we are still working on simplifying this process, and will release this feature soon. Please keep tuned!","title":null,"type":"comment","url":null}],"created_at":"2024-04-01T22:50:56.000Z","created_at_i":1712011856,"id":39900324,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"If only they supported local LLMs out of the box. I have a very specific use case buy it needs to run locally offline only.\nAny suggestions&#x2F;recommendations from fellow HN users are more than welcomed :-)","title":null,"type":"comment","url":null},{"author":"mpeg","children":[{"author":"shekhar101","children":[{"author":"mpeg","children":[{"author":"shekhar101","children":[],"created_at":"2024-04-02T01:44:00.000Z","created_at_i":1712022240,"id":39901507,"options":[],"parent_id":39901031,"points":null,"story_id":39896923,"text":"Thank you very much!","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T00:23:39.000Z","created_at_i":1712017419,"id":39901031,"options":[],"parent_id":39900759,"points":null,"story_id":39896923,"text":"it&#x27;s <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;InfiniFlow&#x2F;deepdoc\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;InfiniFlow&#x2F;deepdoc</a> and the code for usage is in <a href=\"https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;blob&#x2F;main&#x2F;deepdoc&#x2F;README.md\">https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;blob&#x2F;main&#x2F;deepdoc&#x2F;READ...</a> \u2013 it took me a bit of trial and error to get it working<p>It seems to be a YOLOv8 fine-tune, I only did a couple tests but results were decent. Another model that is supposed to be fine tuned for borderless is <a href=\"https:&#x2F;&#x2F;huggingface.co&#x2F;keremberke&#x2F;yolov8m-table-extraction\" rel=\"nofollow\">https:&#x2F;&#x2F;huggingface.co&#x2F;keremberke&#x2F;yolov8m-table-extraction</a> but I haven&#x27;t had great results myself with it, but maybe worth a try for you.","title":null,"type":"comment","url":null},{"author":"thegeomaster","children":[],"created_at":"2024-04-02T22:57:20.000Z","created_at_i":1712098640,"id":39911760,"options":[],"parent_id":39900759,"points":null,"story_id":39896923,"text":"Here&#x27;s a quick test to run: if you have Windows and MS Office, File-&gt;Open your PDF and report the results. You might be surprised at the layout extraction quality.","title":null,"type":"comment","url":null}],"created_at":"2024-04-01T23:44:46.000Z","created_at_i":1712015086,"id":39900759,"options":[],"parent_id":39900338,"points":null,"story_id":39896923,"text":"What&#x27;s the name of the layout recorgniser model? I did not have a good experience extracting layout from tables, especially those without column boundaries (space instead of lines to demarcate boundaries)","title":null,"type":"comment","url":null},{"author":"vissidarte_choi","children":[],"created_at":"2024-04-02T05:07:11.000Z","created_at_i":1712034431,"id":39902572,"options":[],"parent_id":39900338,"points":null,"story_id":39896923,"text":"This is because PDF has so many different versions. A third-party tools like pdfplumber won&#x27;t fit it all. For example, using pdfplumber to parse some PDFs will cause the system to raise exceptions. Sometimes fitz works in situations where pdfplumber won&#x27;t handle well. It looks a bit complicated, but RAGFlow is using multiple parsing tools to handle different types of PDFs.","title":null,"type":"comment","url":null}],"created_at":"2024-04-01T22:52:13.000Z","created_at_i":1712011933,"id":39900338,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"Took me some time to figure out how to run it, but the layout recogniser model hosted on huggingface is pretty good!<p>It correctly identifies tables that even paid models like the AWS Textract Document Analysis API fails to \u2013 for instance tables with one column which often confuse AWS even if they have a clear header and are labelled &quot;Table&quot; in the text.<p>I would however love to know broadly what kind of document it was trained on, as my results could be pure luck, hard to say without a proper benchmark<p>Very nice layout recognition, although I can&#x27;t quite comment on the RAG performance itself \u2013 I think some of the architecture decisions are odd, it mixes a bunch of different PDF parsers for example which will all result in different quality and it&#x27;s not clear to me which one it defaults to as it seems to be different in different places in the code (the simple parser defaults to pypdf2 which is not a great option)","title":null,"type":"comment","url":null},{"author":"cetra3","children":[{"author":"viraptor","children":[],"created_at":"2024-04-01T23:29:49.000Z","created_at_i":1712014189,"id":39900621,"options":[],"parent_id":39900598,"points":null,"story_id":39896923,"text":"It only depends on the interface. There&#x27;s a lot of projects which present the openai interface to whatever you want.","title":null,"type":"comment","url":null},{"author":"vissidarte_choi","children":[{"author":"_akhe","children":[],"created_at":"2024-04-03T19:49:41.000Z","created_at_i":1712173781,"id":39922180,"options":[],"parent_id":39902443,"points":null,"story_id":39896923,"text":"Just link them to <a href=\"https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;blob&#x2F;main&#x2F;rag&#x2F;llm&#x2F;chat_model.py#L132\">https:&#x2F;&#x2F;github.com&#x2F;infiniflow&#x2F;ragflow&#x2F;blob&#x2F;main&#x2F;rag&#x2F;llm&#x2F;chat...</a> :)","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T04:40:28.000Z","created_at_i":1712032828,"id":39902443,"options":[],"parent_id":39900598,"points":null,"story_id":39896923,"text":"Not quite certain about your meaning. Could you be more specific? RAGFlow does not have its own LLM model or souce code. RAGFlow supports API calling from third-party large language model providers, as well as local deployment of these large models. RAGFlow has open-sourced these two parts of codes already.","title":null,"type":"comment","url":null}],"created_at":"2024-04-01T23:26:46.000Z","created_at_i":1712014006,"id":39900598,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"I don&#x27;t consider this strictly open source if components it depends on (I.e, the LLM) is closed source. I&#x27;ve seen a lot of these Fauxpen source style projects around","title":null,"type":"comment","url":null},{"author":"yding","children":[],"created_at":"2024-04-02T01:20:26.000Z","created_at_i":1712020826,"id":39901356,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"This is really cool! Starred and look forward to seeing how this develops further!","title":null,"type":"comment","url":null},{"author":"zzleeper","children":[{"author":"trenchgun","children":[{"author":"jtr101","children":[],"created_at":"2024-04-02T16:51:37.000Z","created_at_i":1712076697,"id":39907943,"options":[],"parent_id":39903459,"points":null,"story_id":39896923,"text":"Love this counterpoint to &quot;OSS means I can get everybody else to do work for me for free&quot; =&gt; &quot;allows you to do the work yourself and share buddy&quot;\nPR or it didn&#x27;t happen","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T08:20:22.000Z","created_at_i":1712046022,"id":39903459,"options":[],"parent_id":39901380,"points":null,"story_id":39896923,"text":"It is open-source though. Just rip it off and make that PDF() class.","title":null,"type":"comment","url":null},{"author":"demilich","children":[],"created_at":"2024-04-02T09:28:53.000Z","created_at_i":1712050133,"id":39903810,"options":[],"parent_id":39901380,"points":null,"story_id":39896923,"text":"Each project has its own detailed requirements and scenarios, and we cannot demand that each project use same library to implement similar functions","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T01:24:09.000Z","created_at_i":1712021049,"id":39901380,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"I&#x27;m partly sad at the approach this and other engines take:  reimplement each part (PDF parser, etc etc) in a way where they are pretty much useless except in their specific engine.<p>If instead we had a PDF() class that did what RAGFlow is doing (dealing with all the different trade-offs of the different python PDF engines such as pdfplumber), then we could easily adapt it and improve it, and it can be useful for other projects as well.","title":null,"type":"comment","url":null},{"author":"bgun","children":[{"author":"scrollaway","children":[{"author":"bgun","children":[{"author":"kaliqt","children":[],"created_at":"2024-04-02T03:25:03.000Z","created_at_i":1712028303,"id":39902080,"options":[],"parent_id":39901725,"points":null,"story_id":39896923,"text":"As a native English speaker I would not dock points from this name for the gender related points as that doesn&#x27;t even cross my mind, it&#x27;s not relevant.<p>Keep gender out of engineering.","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T02:18:45.000Z","created_at_i":1712024325,"id":39901725,"options":[],"parent_id":39901601,"points":null,"story_id":39896923,"text":"I didn&#x27;t point it out because it was funny-ha-ha, but because it&#x27;s a teachable example of a linguistic mis-step that could be easily resolved by diversifying your reviewers - in this case, since you mention it, perhaps a native English speaker?<p>I&#x27;m not even suggesting that this project needs a new name - although I think if they were naming a consumer-facing product or company, someone in marketing would push back on the name almost immediately.","title":null,"type":"comment","url":null},{"author":"kadoban","children":[],"created_at":"2024-04-02T03:08:19.000Z","created_at_i":1712027299,"id":39902001,"options":[],"parent_id":39901601,"points":null,"story_id":39896923,"text":"&gt; Also you might find this normal but I cannot imagine my female colleagues being amused at me calling one of them up asking to \u201cconsult\u201d about a name solely because she\u2019s a woman.<p>I&#x27;d certainly hope it wouldn&#x27;t be &quot;solely&quot; because she&#x27;s a woman. There&#x27;s no people you&#x27;d be interested in input from that happen to be women?<p>A project name is something you&#x27;d want to throw around to a few people (ideally with different perspectives) and make sure it conveys what you intend and has the right tone and such.","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T02:00:33.000Z","created_at_i":1712023233,"id":39901601,"options":[],"parent_id":39901514,"points":null,"story_id":39896923,"text":"The authors are not native English speakers (and what you\u2019re seeing in the name is not something most non native English speakers would spot).<p>I understand it\u2019s funny to point and laugh \u201cha ha dumb dudes didn\u2019t think about something female related because they\u2019re dudes\u201d, but it\u2019s worth remembering that not everyone is from the anglosphere.<p>Also you might find this normal but I cannot imagine my female colleagues being amused at me calling one of them up asking to \u201cconsult\u201d about a name solely because she\u2019s a woman.","title":null,"type":"comment","url":null},{"author":"hitekker","children":[{"author":"bgun","children":[],"created_at":"2024-04-02T02:31:32.000Z","created_at_i":1712025092,"id":39901812,"options":[],"parent_id":39901737,"points":null,"story_id":39896923,"text":"Not sure if sarcasm but &#x27;about as odd as &quot;Nintendo Wii&quot;&#x27; kind of makes my point:<p><a href=\"https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20130623080716&#x2F;http:&#x2F;&#x2F;www.forbes.com&#x2F;2006&#x2F;04&#x2F;28&#x2F;nintendo-wii-console-cx_po_0428autofacescan08.html\" rel=\"nofollow\">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20130623080716&#x2F;http:&#x2F;&#x2F;www.forbes...</a><p><a href=\"http:&#x2F;&#x2F;news.bbc.co.uk&#x2F;2&#x2F;hi&#x2F;technology&#x2F;4953650.stm\" rel=\"nofollow\">http:&#x2F;&#x2F;news.bbc.co.uk&#x2F;2&#x2F;hi&#x2F;technology&#x2F;4953650.stm</a><p><a href=\"https:&#x2F;&#x2F;www.cbc.ca&#x2F;news&#x2F;entertainment&#x2F;why-wii-ask-nintendo-1.612291\" rel=\"nofollow\">https:&#x2F;&#x2F;www.cbc.ca&#x2F;news&#x2F;entertainment&#x2F;why-wii-ask-nintendo-1...</a>","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T02:20:25.000Z","created_at_i":1712024425,"id":39901737,"options":[],"parent_id":39901514,"points":null,"story_id":39896923,"text":"The assumption that author didn&#x27;t talk to a woman sounds pretty knee-jerk to me. Can you explain what you mean?<p>&quot;On the rag&quot; is a semi-common slang for menstruation, but &quot;rag&quot; has way more meanings than just menstrual pads. IMO, as a marketing term, &quot;RAGflow&quot; sounds about as odd as &quot;Nintendo Wii&quot;.","title":null,"type":"comment","url":null},{"author":"nmfisher","children":[],"created_at":"2024-04-02T03:00:01.000Z","created_at_i":1712026801,"id":39901956,"options":[],"parent_id":39901514,"points":null,"story_id":39896923,"text":"People said the same thing about the iPad.","title":null,"type":"comment","url":null},{"author":"yingfeng","children":[],"created_at":"2024-04-02T04:40:40.000Z","created_at_i":1712032840,"id":39902447,"options":[],"parent_id":39901514,"points":null,"story_id":39896923,"text":"Hi, buddy, I&#x27;m sorry to make you feel like we don&#x27;t take women&#x27;s voices into account. We came up with this name because RAG is already a consensus, as an acronym for retrieval augmented generation, RAG is used in many places, and we think it would actually be a standard for LLM oriented B-side scenarios, so that&#x27;s why we adopted this name.","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T01:45:24.000Z","created_at_i":1712022324,"id":39901514,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"I know this is just an open source project but it\u2019s a good example of why you might want to consult a woman before naming things.","title":null,"type":"comment","url":null},{"author":"constantinum","children":[{"author":"TuringNYC","children":[],"created_at":"2024-04-02T04:29:26.000Z","created_at_i":1712032166,"id":39902392,"options":[],"parent_id":39901877,"points":null,"story_id":39896923,"text":"Do you have any recommendations for extraction from Powerpoint documents? Those seem like the worst since each of the layouts tend to be unique (unlike the AMEX Statements in the example)","title":null,"type":"comment","url":null},{"author":"demilich","children":[],"created_at":"2024-04-02T09:30:13.000Z","created_at_i":1712050213,"id":39903817,"options":[],"parent_id":39901877,"points":null,"story_id":39896923,"text":"Cool! This is really helpful.","title":null,"type":"comment","url":null},{"author":"yingfeng","children":[],"created_at":"2024-04-02T09:48:22.000Z","created_at_i":1712051302,"id":39903903,"options":[],"parent_id":39901877,"points":null,"story_id":39896923,"text":"Actually we&#x27;ve tried almost all lof existing open source models for document processing, and none of them performs well for complex documents, especially those having complicated tables, such as tables cells without borders, cells need to be combined,...,etc.  Although adopting LLMs to perform such document understanding tasks is more scalable, it requires much more data and computation power to achieve similar results. That&#x27;s why we design such models start from scratch.","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T02:43:00.000Z","created_at_i":1712025780,"id":39901877,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"Document processing is getting better and better with new tools leveraging LLMs. \nIf anyone is interested in exploring this space, try another similar tool LLMWhisperer (<a href=\"https:&#x2F;&#x2F;llmwhisperer.unstract.com&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;llmwhisperer.unstract.com&#x2F;</a>). It is a part of Unstract, an open-source document processing tool (<a href=\"https:&#x2F;&#x2F;github.com&#x2F;Zipstack&#x2F;unstract\">https:&#x2F;&#x2F;github.com&#x2F;Zipstack&#x2F;unstract</a>)","title":null,"type":"comment","url":null},{"author":"demilich","children":[],"created_at":"2024-04-02T07:11:14.000Z","created_at_i":1712041874,"id":39903108,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"The recognition of file layout to parse file content is indeed a very innovative idea. Compared to many open-source projects we have seen before, this provides us with new ideas for using RAG to solve problems in the future. I hope the author of this project will continue to update it, so that more people can benefit from it.","title":null,"type":"comment","url":null},{"author":"forrest2","children":[{"author":"yingfeng","children":[],"created_at":"2024-04-02T08:08:27.000Z","created_at_i":1712045307,"id":39903392,"options":[],"parent_id":39903304,"points":null,"story_id":39896923,"text":"Thanks for your nice suggestion. We train the model using YOLO, but during inference, the model is converted into ONNX and we use ONNXRuntime for the model inference. As a result, YOLO itself is not included in the software package. \nWe will open the training code in the repo soon.","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T07:52:42.000Z","created_at_i":1712044362,"id":39903304,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"A lot of the yolo stuff from ultralytics is AGPL3 fyi. Recommend caution depending on what code or models &#x2F; model lineage are used","title":null,"type":"comment","url":null},{"author":"StrayHardball","children":[{"author":"zingiscool","children":[],"created_at":"2024-04-02T08:38:22.000Z","created_at_i":1712047102,"id":39903536,"options":[],"parent_id":39903439,"points":null,"story_id":39896923,"text":"Hello friend,<p>I would like to recommend the FigJam feature in Figma to you, a common tool used by high-tech companies for creating Information Architecture (IA). Our company always adheres to the tool usage standards of high-tech companies.<p>For guidance on how to use the FigJam feature in Figma, you can refer to this link: <a href=\"https:&#x2F;&#x2F;youtu.be&#x2F;axDzyLEfYgU?si=V6tqO_tEUKYuLxrL\" rel=\"nofollow\">https:&#x2F;&#x2F;youtu.be&#x2F;axDzyLEfYgU?si=V6tqO_tEUKYuLxrL</a> (or search for FigJam on YouTube).<p>Here are three quick tips on how to efficiently create Information Architecture\uff08Also called System Architecture\uff09:<p>Start with the Sketch method by drawing your product&#x27;s Information architecture on paper. Communicate with your PM, design, and development team to validate its effectiveness.<p>Open FigJam in Figma to turn it into an electronic version. This step is crucial as design and development will follow the version in FigJam for collaborative work.<p>When creating your Information Architecture (IA), first identify the core functionalities of your product, represented by one color. Next, determine what sub-functions each core functionality can be divided into, marked by a second color. Finally, decide what detailed functionalities compose each sub-function, indicated by a third color. Continue in this manner until you complete the entire IA construction.<p>It\u2019s important to note that perfecting the IA is not a linear process; it will go through multiple iterations and modifications. Every great product undergoes this process. Also, initially, you can focus on creating an IA for just one core functional module (usually the innovative feature with the highest user pain point) without defining the entire scope. In essence, flexibly establishing the IA to achieve company goals is the primary task.","title":null,"type":"comment","url":null}],"created_at":"2024-04-02T08:16:18.000Z","created_at_i":1712045778,"id":39903439,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"Does anyone know what software was used to create the &quot;System Architecture&quot; diagram?","title":null,"type":"comment","url":null},{"author":"gardenfelder","children":[],"created_at":"2024-04-02T15:00:55.000Z","created_at_i":1712070055,"id":39906512,"options":[],"parent_id":39896923,"points":null,"story_id":39896923,"text":"I am curious: if I pass a pdf file such as a research report or an open source text book, what will be the result.  Does it create a knowledge graph? triples? Thanks in advance for comments.","title":null,"type":"comment","url":null}],"created_at":"2024-04-01T17:50:54.000Z","created_at_i":1711993854,"id":39896923,"options":[],"parent_id":null,"points":230,"story_id":39896923,"text":null,"title":"RAGFlow is an open-source RAG engine based on OCR and document parsing","type":"story","url":"https://github.com/infiniflow/ragflow"}
