{"author":"saeedesmaili","children":[{"author":"dogline","children":[],"created_at":"2023-07-06T15:38:37.000Z","created_at_i":1688657917,"id":36617682,"options":[],"parent_id":36616799,"points":null,"story_id":36616799,"text":"I hadn&#x27;t seen the unstructured python library before.  Seems handy to parse personal text, like the author is doing.","title":null,"type":"comment","url":null},{"author":"yuppiepuppie","children":[{"author":"saeedesmaili","children":[],"created_at":"2023-07-06T16:00:20.000Z","created_at_i":1688659220,"id":36618036,"options":[],"parent_id":36617990,"points":null,"story_id":36616799,"text":"Author here. That&#x27;s a good point. I&#x27;ll add output examples.","title":null,"type":"comment","url":null}],"created_at":"2023-07-06T15:57:57.000Z","created_at_i":1688659077,"id":36617990,"options":[],"parent_id":36616799,"points":null,"story_id":36616799,"text":"It would help this article\u2019s quality if the author had included an output example from each code snippet. As a reader I\u2019m left to imagine what the output looks like.","title":null,"type":"comment","url":null},{"author":"agadius","children":[{"author":"convivialdingo","children":[{"author":"saeedesmaili","children":[{"author":"convivialdingo","children":[{"author":"saeedesmaili","children":[{"author":"icegreentea2","children":[{"author":"saeedesmaili","children":[{"author":"icegreentea2","children":[],"created_at":"2023-07-06T22:37:43.000Z","created_at_i":1688683063,"id":36624214,"options":[],"parent_id":36622564,"points":null,"story_id":36616799,"text":"Actually, I just realized that I had provided a &#x27;one-off&#x27; hack to a similarish situation here: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;python-openxml&#x2F;python-docx&#x2F;issues&#x2F;1123#issuecomment-1198593723\">https:&#x2F;&#x2F;github.com&#x2F;python-openxml&#x2F;python-docx&#x2F;issues&#x2F;1123#is...</a><p>Replace the `qn(&quot;w:ins&quot;)` in the example with `qn(&quot;w:hyperlink&quot;)` and that should hopefully work?","title":null,"type":"comment","url":null}],"created_at":"2023-07-06T20:35:29.000Z","created_at_i":1688675729,"id":36622564,"options":[],"parent_id":36622326,"points":null,"story_id":36616799,"text":"This sounds amazing! Thanks for sharing it, I will try it to see if I can replace it with the main python-docx. For my use case it suffices to have full text of each paragraph (even if it includes a hyperlink) and heading but also be able to have each of them separated when needed.","title":null,"type":"comment","url":null},{"author":"convivialdingo","children":[],"created_at":"2023-07-06T22:11:36.000Z","created_at_i":1688681496,"id":36623886,"options":[],"parent_id":36622326,"points":null,"story_id":36616799,"text":"Hey, that&#x27;s fantastic.  I&#x27;ll definitely check that out.","title":null,"type":"comment","url":null}],"created_at":"2023-07-06T20:24:03.000Z","created_at_i":1688675043,"id":36622326,"options":[],"parent_id":36621124,"points":null,"story_id":36616799,"text":"Heh, I got a bit into hacking on python-docx last year (the original author seems to be focusing on other things than python-docx now) - I have a fork&#x2F;branch where I tried to more properly implement external hyperlink functionality (<a href=\"https:&#x2F;&#x2F;github.com&#x2F;icegreentea&#x2F;python-docx&#x2F;pull&#x2F;7\">https:&#x2F;&#x2F;github.com&#x2F;icegreentea&#x2F;python-docx&#x2F;pull&#x2F;7</a>)<p>I realize now staring at this, that I might have broken API a little. You can&#x27;t do &quot;text = paragraph.text&quot; anymore, but you can do &quot;text = &#x27;&#x27;.join([run.text for run in paragraph.runs])&quot; instead.<p>If you&#x27;re curious at all why it breaks, it&#x27;s because in the OOXML spec paragraphs are made up of a ordered list of runs or hyperlinks (and hyperlinks can then contain additional runs). The master branch just implements paragraphs as ordered list of runs (and ignores all hyperlinks).","title":null,"type":"comment","url":null}],"created_at":"2023-07-06T19:03:40.000Z","created_at_i":1688670220,"id":36621124,"options":[],"parent_id":36620904,"points":null,"story_id":36616799,"text":"I tried python-docx with a bunch of docx files (downloaded from Google Docs). It returns empty strings for hyperlinks and I couldn&#x27;t manage to fix this. So if there is a sentence like &quot;This is an important link to another doc or url.&quot; and the &quot;link&quot; is a hyperlink, python-docx returns &quot;This is an important  to another doc or url.&quot;","title":null,"type":"comment","url":null}],"created_at":"2023-07-06T18:50:07.000Z","created_at_i":1688669407,"id":36620904,"options":[],"parent_id":36618948,"points":null,"story_id":36616799,"text":"I&#x27;ve had good luck with python-docx for reading word documents (typically specifications).  Tables are supported - but it&#x27;s not obvious where the table comes from in the document and I had to come up with a hack way to read image captions.<p>PDF has been hit or miss, but pypdf has improved in the last couple of years.  Depending on the document you&#x27;ll sometimes get   random    spaces or nospacesatall.","title":null,"type":"comment","url":null}],"created_at":"2023-07-06T16:52:36.000Z","created_at_i":1688662356,"id":36618948,"options":[],"parent_id":36618877,"points":null,"story_id":36616799,"text":"Do you have any suggestions for Python libraries (other than what&#x27;s mentioned in the post)?","title":null,"type":"comment","url":null}],"created_at":"2023-07-06T16:48:58.000Z","created_at_i":1688662138,"id":36618877,"options":[],"parent_id":36618056,"points":null,"story_id":36616799,"text":"I second this suggestion.  I tested numerous Python tools to extract text - nothing matches Tika for general extraction of just about any data format.<p>However - if you can expect a certain format beforehand - then Python is better since you can extract higher-quality data (tables, lists) with the appropriate tool.","title":null,"type":"comment","url":null},{"author":"mcswell","children":[],"created_at":"2023-07-06T23:08:46.000Z","created_at_i":1688684926,"id":36624564,"options":[],"parent_id":36618056,"points":null,"story_id":36616799,"text":"Tika can be used as a library in Python: <a href=\"https:&#x2F;&#x2F;pypi.org&#x2F;project&#x2F;tika&#x2F;\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;pypi.org&#x2F;project&#x2F;tika&#x2F;</a>","title":null,"type":"comment","url":null},{"author":"ramraj07","children":[{"author":"mbwgh","children":[{"author":"mcswell","children":[],"created_at":"2023-07-07T15:49:43.000Z","created_at_i":1688744983,"id":36633533,"options":[],"parent_id":36629257,"points":null,"story_id":36616799,"text":"There&#x27;s a book called &quot;Tika in action&quot; which I found useful.","title":null,"type":"comment","url":null}],"created_at":"2023-07-07T09:28:31.000Z","created_at_i":1688722111,"id":36629257,"options":[],"parent_id":36627438,"points":null,"story_id":36616799,"text":"I second this, there is absolutely no easily discoverable entry point to the documentation.\nIn the end if you want to get a feeling of what this is you search for &quot;tika tutorial&quot; and get a rough idea via (in my case) some medium article I guess.","title":null,"type":"comment","url":null},{"author":"mcswell","children":[],"created_at":"2023-07-07T15:48:12.000Z","created_at_i":1688744892,"id":36633523,"options":[],"parent_id":36627438,"points":null,"story_id":36616799,"text":"I can&#x27;t speak to the Apache documentation, but I once had the task of extracting plain text from many different document formats: Word, spreadsheets, PDFs, the EXIF information in JPEGs, and so on for a long list.  I had written code with calls to extractor libraries for several of these formats, when I can across tika.  Out when my if..then..elif..elif..elif.. code, to be replaced with a single (Python) call to tika.<p>I can&#x27;t answer your question about pandas, though.","title":null,"type":"comment","url":null}],"created_at":"2023-07-07T05:11:10.000Z","created_at_i":1688706670,"id":36627438,"options":[],"parent_id":36618056,"points":null,"story_id":36616799,"text":"As is customary for all of Apache, I have no clue what I\u2019m looking at after trying to read through the links in that page for ten minutes. Like who is this tool for? When should I use this vs any other competing tools? No clue. I suppose it can read documents of any type and give it out as a dictionary? Why would I use this vs pandas?","title":null,"type":"comment","url":null},{"author":"jghn","children":[],"created_at":"2023-07-07T05:49:28.000Z","created_at_i":1688708968,"id":36627711,"options":[],"parent_id":36618056,"points":null,"story_id":36616799,"text":"I found myself today trying to parse a TSV and substituting a few fields with a different value, then writing the new file out.<p>Something that perl would excel at, although I used Python. Because Perl isn&#x27;t as maintainable as Python<p>I was intrigued by this comment. A JVM solution would also be viable in my tech stack. Would Tika be easier than line processing compiled regexes in Python? I tried looking at the Usage examples but it wasn&#x27;t clear.","title":null,"type":"comment","url":null}],"created_at":"2023-07-06T16:01:27.000Z","created_at_i":1688659287,"id":36618056,"options":[],"parent_id":36616799,"points":null,"story_id":36616799,"text":"If you accept running Java, the Apache Tika is extremely good at parsing content (<a href=\"https:&#x2F;&#x2F;tika.apache.org&#x2F;\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;tika.apache.org&#x2F;</a>)","title":null,"type":"comment","url":null},{"author":"CShorten","children":[],"created_at":"2023-07-06T16:10:10.000Z","created_at_i":1688659810,"id":36618184,"options":[],"parent_id":36616799,"points":null,"story_id":36616799,"text":"Awesome!","title":null,"type":"comment","url":null},{"author":"oersted","children":[],"created_at":"2023-07-06T16:34:20.000Z","created_at_i":1688661260,"id":36618587,"options":[],"parent_id":36616799,"points":null,"story_id":36616799,"text":"I have been using it extensively during the last few weeks. I&#x27;ve very thankful for such a clean and practical API, and I think it will become the central solution for ingesting heterogeneous text in the Python ecosystem.<p>However, I&#x27;m afraid it is not there yet. Other libraries like PDFMiner give higher quality outputs and specialized libraries like Camelot are still needed to extract tables as reasonably well formatted text. It also needs a lot of extra tooling for web scraping. Sure it can read plain HTML from a URL, but it cannot run JavaScript, or control things like User Agent. It could be argued that such features are not within the scope, but it is rather bothersome for a library that presents a magic `partition` function for most standard text sources.<p>I&#x27;m sure it will get there soon though. It shouldn&#x27;t be hard to integrate with state-of-the-art parsers and tooling, and the simple API undoubtedly brings a lot of peace of mind.","title":null,"type":"comment","url":null},{"author":"froggychairs","children":[],"created_at":"2023-07-06T18:32:17.000Z","created_at_i":1688668337,"id":36620649,"options":[],"parent_id":36616799,"points":null,"story_id":36616799,"text":"API looks very clean :) I\u2019ve also been avoiding LangChain since it just seems too big for my tastes. Will give this a shot","title":null,"type":"comment","url":null},{"author":"raphman","children":[],"created_at":"2023-07-07T02:59:44.000Z","created_at_i":1688698784,"id":36626533,"options":[],"parent_id":36616799,"points":null,"story_id":36616799,"text":"Another way to parse Markdown, HTML, or docx files would be pandoc [1]:<p><pre><code>  pandoc --to json file.docx\n</code></pre>\nor in Python:<p><pre><code>  import json\n  from sh import pandoc\n  doc = json.loads(  pandoc(&quot;file.docx&quot;, to=&quot;json&quot;).stdout  )\n</code></pre>\nExample output (reformatted slightly to reduce number of lines:<p><pre><code>  {&#x27;pandoc-api-version&#x27;: [1, 22, 2],\n   &#x27;meta&#x27;: {&#x27;title&#x27;: {&#x27;t&#x27;: &#x27;MetaInlines&#x27;,\n     &#x27;c&#x27;: [{&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;The&#x27;}, {&#x27;t&#x27;: &#x27;Space&#x27;}, {&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;Title&#x27;}]}},\n   &#x27;blocks&#x27;: [{&#x27;t&#x27;: &#x27;Header&#x27;,\n     &#x27;c&#x27;: [1,\n      [&#x27;first-chapter&#x27;, [], []],\n      [{&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;First&#x27;}, {&#x27;t&#x27;: &#x27;Space&#x27;}, {&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;Chapter&#x27;}]]},\n    {&#x27;t&#x27;: &#x27;Para&#x27;,\n     &#x27;c&#x27;: [{&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;I&#x27;}, {&#x27;t&#x27;: &#x27;Space&#x27;}, {&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;like&#x27;}, {&#x27;t&#x27;: &#x27;Space&#x27;},\n      {&#x27;t&#x27;: &#x27;Emph&#x27;, &#x27;c&#x27;: [{&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;cursive&#x27;}]}, {&#x27;t&#x27;: &#x27;Space&#x27;}, {&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;or&#x27;}, \n      {&#x27;t&#x27;: &#x27;Space&#x27;}, {&#x27;t&#x27;: &#x27;Strong&#x27;, &#x27;c&#x27;: [{&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;bold&#x27;}]}, {&#x27;t&#x27;: &#x27;Space&#x27;},\n      {&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;text.&#x27;}]},\n    {&#x27;t&#x27;: &#x27;Para&#x27;,\n     &#x27;c&#x27;: [{&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;Here&#x27;}, {&#x27;t&#x27;: &#x27;Space&#x27;}, {&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;is&#x27;}, {&#x27;t&#x27;: &#x27;Space&#x27;},\n      {&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;a&#x27;}, {&#x27;t&#x27;: &#x27;Space&#x27;}, {&#x27;t&#x27;: &#x27;Link&#x27;,\n       &#x27;c&#x27;: [[&#x27;&#x27;, [], []], [{&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;link&#x27;}], [&#x27;https:&#x2F;&#x2F;ix.de&#x2F;&#x27;, &#x27;&#x27;]]},\n      {&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;.&#x27;}]},\n    {&#x27;t&#x27;: &#x27;BulletList&#x27;,\n     &#x27;c&#x27;: [[{&#x27;t&#x27;: &#x27;Para&#x27;, &#x27;c&#x27;: [{&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;Item&#x27;}, {&#x27;t&#x27;: &#x27;Space&#x27;}, &#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;1&#x27;}]}],\n      [{&#x27;t&#x27;: &#x27;Para&#x27;, &#x27;c&#x27;: [{&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;Item&#x27;}, {&#x27;t&#x27;: &#x27;Space&#x27;}, {&#x27;t&#x27;: &#x27;Str&#x27;, &#x27;c&#x27;: &#x27;2&#x27;}]}]]}]}\n</code></pre>\n[1] <a href=\"https:&#x2F;&#x2F;pandoc.org&#x2F;\" rel=\"nofollow noreferrer\">https:&#x2F;&#x2F;pandoc.org&#x2F;</a>","title":null,"type":"comment","url":null}],"created_at":"2023-07-06T14:47:58.000Z","created_at_i":1688654878,"id":36616799,"options":[],"parent_id":null,"points":141,"story_id":36616799,"text":null,"title":"Demystifying Text Data with the Unstructured Python Library","type":"story","url":"https://saeedesmaili.com/demystifying-text-data-with-the-unstructured-python-library/"}
