{"author":"vasa_","children":[{"author":"dandan7","children":[],"created_at":"2025-02-13T10:50:53.000Z","created_at_i":1739443853,"id":43034688,"options":[],"parent_id":43031915,"points":null,"story_id":43031915,"text":"Good luck guys looks great","title":null,"type":"comment","url":null}],"created_at":"2025-02-13T01:59:26.000Z","created_at_i":1739411966,"id":43031915,"options":[],"parent_id":null,"points":6,"story_id":43031915,"text":"Hey there HN! We\u2019re Vasilije, Boris, and Laszlo, and we\u2019re excited to introduce cognee, an open-source Python library that approaches building evolving semantic memory using knowledge graphs + data pipelines<p>Before we built cognee, Vasilije(B Economics and Clinical Psychology) worked at a few unicorns (Omio, Zalando, Taxfix), while Boris managed large-scale applications in production at Pera and StuDocu. Laszlo joined after getting his PhD in Graph Theory at the University of Szeged.<p>Using LLMs to connect to large datasets (RAG) has been popularized and has shown great promise. Unfortunately, this approach doesn\u2019t live up to the hype.<p>Let\u2019s assume we want to load a large repository from GitHub to a vector store.\nConnectingfiles in larger systems with RAG would fail because a fixed RAG limit is too constraining in longer dependency chains. While we need results that are aware of the context of the whole repository, RAG\u2019s similarity-based retrieval does not capture the full context of interdependent files spread across the repository.<p>This approach allows cognee to retrieve all relevant and correct context at inference time. For example, if `function A` in one file calls `function B` in another file, which calls `function C` in a third file, all code and summaries that further explain their position and purpose in that chain are served as context. As a result, the system has complete visibility into how different code parts work together within the repo.<p>Last year, Microsoft took a leap published GraphRAG - i.e. RAG with Knowledge Graphs. We think it is the right direction.\nOur initial ideas were similar to this paper and this got some attention on Twitter (<a href=\"https:&#x2F;&#x2F;x.com&#x2F;tricalt&#x2F;status&#x2F;1722216426709365024\" rel=\"nofollow\">https:&#x2F;&#x2F;x.com&#x2F;tricalt&#x2F;status&#x2F;1722216426709365024</a>)<p>Over time we understood we needed tooling to create dynamically evolving groups of graphs, cross-connected and evaluated together.\nOur tool is named after a process called cognification. We prefer the definition that Vakalo (1978) uses to explain that cognify represents &quot;building a fitting (mental) picture&quot;<p>We believe that agents of tomorrow will require a correct dynamic \u201cmental picture\u201d or context to operate in a rapidly evolving landscape.<p>To address this, we built ECL pipelines, where we do the following:\n- Extract data from various sources using dlt and existing frameworks\n- Cognify - create a graph&#x2F;vector representation of the data\n- Load - store the data in the vector (in this case our partner FalkorDB), graph, and relational stores<p>We can also continuously feed the graph with new information, and when testing this approach we found that on HotpotQA, with human labeling, we achieved 87% answer accuracy (<a href=\"https:&#x2F;&#x2F;docs.cognee.ai&#x2F;evaluations\" rel=\"nofollow\">https:&#x2F;&#x2F;docs.cognee.ai&#x2F;evaluations</a>).<p>To show how the approach works we did an integration with continue.dev and built a codegraph<p>Here is how codegraph was implemented: \nWe&#x27;re explicitly including repository structure details and integrating custom dependency graph versions. Think of it as a more insightful way to understand your codebase&#x27;s architecture.\nBy transforming dependency graphs into knowledge graphs, we&#x27;re creating a quick, graph-based version of tools like tree-sitter. This means faster and more accurate code analysis.\nWe worked on modeling causal relationships within code and enriching them with LLMs. This helps you understand how different parts of your code influence each other.\nWe created graph skeletons in memory which allows us to perform various operations on graphs and power custom retrievers.<p>If you want to integrate cognee into your systems or have a look at codegraph, our GitHub repository is (<a href=\"https:&#x2F;&#x2F;github.com&#x2F;topoteretes&#x2F;cognee\">https:&#x2F;&#x2F;github.com&#x2F;topoteretes&#x2F;cognee</a>)<p>Thank you for reading! We\u2019re definitely early and welcome your ideas and experiences as it relates to agents, graphs, evals, and human+LLM memory.","title":"Show HN: Cognee \u2013 Turn RAG and GraphRAG into custom dynamic semantic memory","type":"story","url":"https://github.com/topoteretes/cognee"}
