{"exhaustive":{"nbHits":false,"typo":false},"exhaustiveNbHits":false,"exhaustiveTypo":false,"hits":[{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"arauhala"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"AI <em>Database</em> Landscape in <em>2026</em>: <em>Vector</em>, ML-in-DB, LLM-Augmented, Predictive"},"url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["database","2026"],"value":"https://aito.ai/blog/the-ai-<em>database</em>-landscape-in-<em>2026</em>-where-does-structured-prediction-fit/"}},"_tags":["story","author_arauhala","story_47847289"],"author":"arauhala","children":[47847297],"created_at":"2026-04-21T11:25:12Z","created_at_i":1776770712,"num_comments":1,"objectID":"47847289","points":2,"story_id":47847289,"title":"AI Database Landscape in 2026: Vector, ML-in-DB, LLM-Augmented, Predictive","updated_at":"2026-04-21T17:56:21Z","url":"https://aito.ai/blog/the-ai-database-landscape-in-2026-where-does-structured-prediction-fit/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"sirily11"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"Application Signal is a platform that collects every YC-backed startup from 2020 to <em>2026</em>, embeds them into a <em>vector</em> <em>database</em>, and visualizes them on an interactive map. It also provide an AI-powered analysis service. Users can generate a detailed report for their startup idea using credits."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Application Signal \u2013 AI That Evaluates Your YC Startup Idea"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://ycreport.rxlab.app"}},"_tags":["story","author_sirily11","story_48993394","show_hn"],"author":"sirily11","children":[48993594],"created_at":"2026-07-21T15:16:20Z","created_at_i":1784646980,"num_comments":1,"objectID":"48993394","points":2,"story_id":48993394,"story_text":"Application Signal is a platform that collects every YC-backed startup from 2020 to 2026, embeds them into a vector database, and visualizes them on an interactive map. It also provide an AI-powered analysis service. Users can generate a detailed report for their startup idea using credits.","title":"Show HN: Application Signal \u2013 AI That Evaluates Your YC Startup Idea","updated_at":"2026-07-22T02:07:58Z","url":"https://ycreport.rxlab.app"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"haensi"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"I\u2019ve been keeping up with the classics (NCF, Wide &amp; Deep, LightGCN), but the field seems to have shifted dramatically in the last 18\u201324 months toward LLM-based reasoning and graph-based retrieval at scale.<p>I\u2019m looking for the &quot;state of the art&quot; in <em>2026</em>. Specifically:<p>LLM4Rec: Beyond just using LLMs for feature engineering\u2014who is doing generative recommendation well?<p>Retrieval vs. Ranking: Any new breakthroughs in the &quot;Two-Tower&quot; paradigm or <em>vector</em> <em>database</em> integration?<p>Real-world Scale: Papers that address the latency/cost trade-offs of these newer, heavier models.<p>What has been the most influential paper you\u2019ve read recently that changed how you think about discovery?"},"title":{"matchLevel":"none","matchedWords":[],"value":"Ask HN: What are the recommender systems papers from 2024-2025?"}},"_tags":["story","author_haensi","story_46692368","ask_hn"],"author":"haensi","children":[46700418],"created_at":"2026-01-20T14:50:24Z","created_at_i":1768920624,"num_comments":1,"objectID":"46692368","points":15,"story_id":46692368,"story_text":"I\u2019ve been keeping up with the classics (NCF, Wide &amp; Deep, LightGCN), but the field seems to have shifted dramatically in the last 18\u201324 months toward LLM-based reasoning and graph-based retrieval at scale.<p>I\u2019m looking for the &quot;state of the art&quot; in 2026. Specifically:<p>LLM4Rec: Beyond just using LLMs for feature engineering\u2014who is doing generative recommendation well?<p>Retrieval vs. Ranking: Any new breakthroughs in the &quot;Two-Tower&quot; paradigm or vector database integration?<p>Real-world Scale: Papers that address the latency&#x2F;cost trade-offs of these newer, heavier models.<p>What has been the most influential paper you\u2019ve read recently that changed how you think about discovery?","title":"Ask HN: What are the recommender systems papers from 2024-2025?","updated_at":"2026-03-05T23:29:02Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"yoloshii"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"Simple Apify actor that scrapes websites and indexes them in Google's new Gemini File Search API (launched Nov 6).<p>The workflow: Scrape \u2192 Clean content \u2192 Upload to Gemini \u2192 Get permanent queryable knowledge base with automatic citations.<p>Technical approach:\n- Intelligent scraper selection from 5 Apify-native scrapers\n- Automatic content cleaning (strips nav/ads/fluff)\n- Uploads to Gemini File Search (persistent storage)\n- Per-page PPE pricing ($0.02 start, $0.0015/page)<p>Potential use cases:\n- Turn documentation into AI chatbots\n- Make company wikis naturally searchable\n- Build RAG apps without managing <em>vector</em> <em>databases</em><p>Basically streamline your Gemini file search RAG ingestion with an Apify scraper run AIO.<p>You can also employ the actor programmatically with agents using Apify's actor mcp.<p>Built this over a weekend for the Apify 1M Actor Challenge. It's my first Apify actor, so curious to hear if the pricing makes sense.<p><i>Note*</i> There is a banned websites to scrape list filter due to the constraints of the challenge. This will be lifted after the challenge ends (Jan 31, <em>2026</em>)."},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Scrape websites into queryable Gemini RAG knowledge bases"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://apify.com/yoloshii/gemini-file-search-builder"}},"_tags":["story","author_yoloshii","story_46260371","show_hn"],"author":"yoloshii","created_at":"2025-12-14T02:46:32Z","created_at_i":1765680392,"num_comments":0,"objectID":"46260371","points":1,"story_id":46260371,"story_text":"Simple Apify actor that scrapes websites and indexes them in Google&#x27;s new Gemini File Search API (launched Nov 6).<p>The workflow: Scrape \u2192 Clean content \u2192 Upload to Gemini \u2192 Get permanent queryable knowledge base with automatic citations.<p>Technical approach:\n- Intelligent scraper selection from 5 Apify-native scrapers\n- Automatic content cleaning (strips nav&#x2F;ads&#x2F;fluff)\n- Uploads to Gemini File Search (persistent storage)\n- Per-page PPE pricing ($0.02 start, $0.0015&#x2F;page)<p>Potential use cases:\n- Turn documentation into AI chatbots\n- Make company wikis naturally searchable\n- Build RAG apps without managing vector databases<p>Basically streamline your Gemini file search RAG ingestion with an Apify scraper run AIO.<p>You can also employ the actor programmatically with agents using Apify&#x27;s actor mcp.<p>Built this over a weekend for the Apify 1M Actor Challenge. It&#x27;s my first Apify actor, so curious to hear if the pricing makes sense.<p><i>Note*</i> There is a banned websites to scrape list filter due to the constraints of the challenge. This will be lifted after the challenge ends (Jan 31, 2026).","title":"Show HN: Scrape websites into queryable Gemini RAG knowledge bases","updated_at":"2026-03-05T23:11:30Z","url":"https://apify.com/yoloshii/gemini-file-search-builder"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"StephSpanjian"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"I shelved a project in March 2025 because it cost $360/month to run.<p>I rebuilt it in January <em>2026</em> for $1.12/month.<p>Same functionality. Better architecture. No <em>vector</em> <em>database</em>.<p>It's a RAG pipeline \u2014 API Gateway \u2192 Lambda \u2192 Amazon Bedrock \u2192 S3 embeddings \u2192 in-memory cosine similarity search. It lives on my site and answers questions about my experience and work history in real time, grounded in structured data I built and curated myself.<p>The interesting parts:\n\u2014 Why I killed OpenSearch after less than a week\n\u2014 How in-memory search outperformed a network call at this scale\n\u2014 The Anthropic/Bedrock access issue I still haven't fully resolved (Llama works fine)\n\u2014 Why I handcrafted every knowledge base chunk instead of automating it"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"RAG on a Budget: How I Replaced a $360/Month OpenSearch Cluster for $1.12/Month"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://stephaniespanjian.com/blog/rag-cost-reduction-replaced-opensearch-s3-in-memory-search"}},"_tags":["comment","author_StephSpanjian","story_47160232"],"author":"StephSpanjian","comment_text":"I shelved a project in March 2025 because it cost $360&#x2F;month to run.<p>I rebuilt it in January 2026 for $1.12&#x2F;month.<p>Same functionality. Better architecture. No vector database.<p>It&#x27;s a RAG pipeline \u2014 API Gateway \u2192 Lambda \u2192 Amazon Bedrock \u2192 S3 embeddings \u2192 in-memory cosine similarity search. It lives on my site and answers questions about my experience and work history in real time, grounded in structured data I built and curated myself.<p>The interesting parts:\n\u2014 Why I killed OpenSearch after less than a week\n\u2014 How in-memory search outperformed a network call at this scale\n\u2014 The Anthropic&#x2F;Bedrock access issue I still haven&#x27;t fully resolved (Llama works fine)\n\u2014 Why I handcrafted every knowledge base chunk instead of automating it","created_at":"2026-02-26T00:38:33Z","created_at_i":1772066313,"objectID":"47160233","parent_id":47160232,"story_id":47160232,"story_title":"RAG on a Budget: How I Replaced a $360/Month OpenSearch Cluster for $1.12/Month","story_url":"https://stephaniespanjian.com/blog/rag-cost-reduction-replaced-opensearch-s3-in-memory-search","updated_at":"2026-03-05T23:37:31Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"westurner"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"Np. Still learning about how to reference agentskills from my coding LLM here. Haven't yet worked with MCP or A2A much. I should port instruction prompts from each project's AGENTS.md to reusable skills in .claude/skills/skillname/SKILL.md ; so that I can reference &quot;/skillname&quot;  in a prompt to reuse that prompt text.<p>Copilot looks for agent skills in .claude/skills/, .github/skills/, or .agents/skills/ .<p>FWIU there are package managers for agent skills?<p>PAKS: \n<a href=\"https://github.com/stakpak/paks\" rel=\"nofollow\">https://github.com/stakpak/paks</a><p>vercel-labs/skills: \n<a href=\"https://github.com/vercel-labs/skills\" rel=\"nofollow\">https://github.com/vercel-labs/skills</a> :<p><pre><code>  export DISABLE_TELEMETRY=1\n  npx skills add vercel-labs/agent-skills\n</code></pre>\n<a href=\"https://skills.sh/docs/cli\" rel=\"nofollow\">https://skills.sh/docs/cli</a><p>A reproducibility, traceability, and auditability concern:<p>How to show the text of the exact version of the skill referenced in each input prompt in each chat log? GH copilot's JSON export format is a fairly complete trace compared to e.g. Gemini Chat.<p>.<p>ElasticSearch (Java) and  Meilisearch (Rust) do at least porter stemming before indexing.<p>MeiliSearch can be used as a <em>vector</em> <em>database</em> for langchain.<p>meilisearch/meilisearch: <a href=\"https://github.com/meilisearch/meilisearch\" rel=\"nofollow\">https://github.com/meilisearch/meilisearch</a> :<p>Meilisearch docs &gt; langchain guide &gt; Performing Similarity Search: \n<a href=\"https://www.meilisearch.com/docs/guides/langchain#performing-similarity-search\" rel=\"nofollow\">https://www.meilisearch.com/docs/guides/langchain#performing...</a><p><pre><code>  from langchain.vectorstores import Meilisearch\n  from langchain.embeddings.openai import OpenAIEmbeddings\n</code></pre>\nmeilisearch/meilisearch-mcp :\n<a href=\"https://github.com/meilisearch/meilisearch-mcp\" rel=\"nofollow\">https://github.com/meilisearch/meilisearch-mcp</a><p>openobserve/openobserve : \n<a href=\"https://github.com/openobserve/openobserve\" rel=\"nofollow\">https://github.com/openobserve/openobserve</a> :<p>&gt; <i>OpenObserve is an open-source observability platform for logs, metrics, traces, and frontend monitoring. A cost-effective alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage costs and single binary deployment.</i> with SQL in Rust<p>OpenObserve (AGPL) docs &gt; MCP: \n<a href=\"https://openobserve.ai/docs/integration/mcp/\" rel=\"nofollow\">https://openobserve.ai/docs/integration/mcp/</a> :<p>&gt; <i>This capability is supported in the Enterprise edition of OpenObserve</i><p><a href=\"https://github.com/openobserve/openobserve/issues/6804\" rel=\"nofollow\">https://github.com/openobserve/openobserve/issues/6804</a><p>openobserve docs &gt; Comparison with Alternatives &gt; \nHow does OpenObserve compare to Elasticsearch: \n<a href=\"https://openobserve.ai/docs/overview/comparison/\" rel=\"nofollow\">https://openobserve.ai/docs/overview/comparison/</a><p>meilisearch blog &gt; &quot;What is GraphRAG: Complete guide [<em>2026</em>]&quot;  <a href=\"https://www.meilisearch.com/blog/graph-rag\" rel=\"nofollow\">https://www.meilisearch.com/blog/graph-rag</a> :<p>&gt; <i>How is GraphRAG different from baseline RAG?</i><p>&gt; <i>GraphRAG represents one of the different types of RAG, and its retrieval process differs from baseline RAG.</i><p>&gt; <i>Baseline RAG is <em>vector</em> search-based, while GraphRAG uses structured relationships to get the end result. GraphRAG can still use <em>vector</em> and full-text search, but relationships drive what gets retrieved.</i><p>/? graphrag sqlite fts :  <a href=\"https://www.google.com/search?q=graphrag+sqlite+fts\" rel=\"nofollow\">https://www.google.com/search?q=graphrag+sqlite+fts</a>"},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: Saga \u2013 A Jira-like project tracker MCP server for AI agents (SQLite)"}},"_tags":["comment","author_westurner","story_47106215"],"author":"westurner","children":[47159534],"comment_text":"Np. Still learning about how to reference agentskills from my coding LLM here. Haven&#x27;t yet worked with MCP or A2A much. I should port instruction prompts from each project&#x27;s AGENTS.md to reusable skills in .claude&#x2F;skills&#x2F;skillname&#x2F;SKILL.md ; so that I can reference &quot;&#x2F;skillname&quot;  in a prompt to reuse that prompt text.<p>Copilot looks for agent skills in .claude&#x2F;skills&#x2F;, .github&#x2F;skills&#x2F;, or .agents&#x2F;skills&#x2F; .<p>FWIU there are package managers for agent skills?<p>PAKS: \n<a href=\"https:&#x2F;&#x2F;github.com&#x2F;stakpak&#x2F;paks\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;stakpak&#x2F;paks</a><p>vercel-labs&#x2F;skills: \n<a href=\"https:&#x2F;&#x2F;github.com&#x2F;vercel-labs&#x2F;skills\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;vercel-labs&#x2F;skills</a> :<p><pre><code>  export DISABLE_TELEMETRY=1\n  npx skills add vercel-labs&#x2F;agent-skills\n</code></pre>\n<a href=\"https:&#x2F;&#x2F;skills.sh&#x2F;docs&#x2F;cli\" rel=\"nofollow\">https:&#x2F;&#x2F;skills.sh&#x2F;docs&#x2F;cli</a><p>A reproducibility, traceability, and auditability concern:<p>How to show the text of the exact version of the skill referenced in each input prompt in each chat log? GH copilot&#x27;s JSON export format is a fairly complete trace compared to e.g. Gemini Chat.<p>.<p>ElasticSearch (Java) and  Meilisearch (Rust) do at least porter stemming before indexing.<p>MeiliSearch can be used as a vector database for langchain.<p>meilisearch&#x2F;meilisearch: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;meilisearch&#x2F;meilisearch\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;meilisearch&#x2F;meilisearch</a> :<p>Meilisearch docs &gt; langchain guide &gt; Performing Similarity Search: \n<a href=\"https:&#x2F;&#x2F;www.meilisearch.com&#x2F;docs&#x2F;guides&#x2F;langchain#performing-similarity-search\" rel=\"nofollow\">https:&#x2F;&#x2F;www.meilisearch.com&#x2F;docs&#x2F;guides&#x2F;langchain#performing...</a><p><pre><code>  from langchain.vectorstores import Meilisearch\n  from langchain.embeddings.openai import OpenAIEmbeddings\n</code></pre>\nmeilisearch&#x2F;meilisearch-mcp :\n<a href=\"https:&#x2F;&#x2F;github.com&#x2F;meilisearch&#x2F;meilisearch-mcp\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;meilisearch&#x2F;meilisearch-mcp</a><p>openobserve&#x2F;openobserve : \n<a href=\"https:&#x2F;&#x2F;github.com&#x2F;openobserve&#x2F;openobserve\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;openobserve&#x2F;openobserve</a> :<p>&gt; <i>OpenObserve is an open-source observability platform for logs, metrics, traces, and frontend monitoring. A cost-effective alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage costs and single binary deployment.</i> with SQL in Rust<p>OpenObserve (AGPL) docs &gt; MCP: \n<a href=\"https:&#x2F;&#x2F;openobserve.ai&#x2F;docs&#x2F;integration&#x2F;mcp&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;openobserve.ai&#x2F;docs&#x2F;integration&#x2F;mcp&#x2F;</a> :<p>&gt; <i>This capability is supported in the Enterprise edition of OpenObserve</i><p><a href=\"https:&#x2F;&#x2F;github.com&#x2F;openobserve&#x2F;openobserve&#x2F;issues&#x2F;6804\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;openobserve&#x2F;openobserve&#x2F;issues&#x2F;6804</a><p>openobserve docs &gt; Comparison with Alternatives &gt; \nHow does OpenObserve compare to Elasticsearch: \n<a href=\"https:&#x2F;&#x2F;openobserve.ai&#x2F;docs&#x2F;overview&#x2F;comparison&#x2F;\" rel=\"nofollow\">https:&#x2F;&#x2F;openobserve.ai&#x2F;docs&#x2F;overview&#x2F;comparison&#x2F;</a><p>meilisearch blog &gt; &quot;What is GraphRAG: Complete guide [2026]&quot;  <a href=\"https:&#x2F;&#x2F;www.meilisearch.com&#x2F;blog&#x2F;graph-rag\" rel=\"nofollow\">https:&#x2F;&#x2F;www.meilisearch.com&#x2F;blog&#x2F;graph-rag</a> :<p>&gt; <i>How is GraphRAG different from baseline RAG?</i><p>&gt; <i>GraphRAG represents one of the different types of RAG, and its retrieval process differs from baseline RAG.</i><p>&gt; <i>Baseline RAG is vector search-based, while GraphRAG uses structured relationships to get the end result. GraphRAG can still use vector and full-text search, but relationships drive what gets retrieved.</i><p>&#x2F;? graphrag sqlite fts :  <a href=\"https:&#x2F;&#x2F;www.google.com&#x2F;search?q=graphrag+sqlite+fts\" rel=\"nofollow\">https:&#x2F;&#x2F;www.google.com&#x2F;search?q=graphrag+sqlite+fts</a>","created_at":"2026-02-25T23:05:14Z","created_at_i":1772060714,"objectID":"47159322","parent_id":47143490,"story_id":47106215,"story_title":"Show HN: Saga \u2013 A Jira-like project tracker MCP server for AI agents (SQLite)","updated_at":"2026-03-05T23:37:28Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"ckarani"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"I built Wax because every RAG solution required either Pinecone/Weaviate in the cloud or ChromaDB/Qdrant running locally. I wanted the SQLite of RAG -- import a library, open a file, query. Except for multimodal content at GPU speed.<p>The architecture that makes this work:\nMetal-accelerated <em>vector</em> search -- Embeddings live directly in unified memory (MTLBuffer). Zero CPU-GPU copy overhead. Adaptive SIMD4/SIMD8 kernels + GPU-side bitonic sort = sub-millisecond search on 10K+ vectors (vs ~100ms CPU). This isn't just &quot;faster&quot; -- it enables interactive search UX that wasn't possible before.<p>Atomic single-file storage (.mv2s) -- Everything in one crash-safe binary: embeddings, BM25 index, metadata, compressed payloads. Dual-header writes with generation counters = kill -9 safe. Sync via iCloud, email it, commit to git. The file format is deterministic -- identical input produces byte-identical output.<p>Query-adaptive hybrid fusion -- Four parallel search lanes (BM25, <em>vector</em>, timeline, structured memory). Lightweight classifier detects intent (&quot;when did I...&quot; \u2192 boost timeline, &quot;find documentation about...&quot; \u2192 boost BM25). Reciprocal Rank Fusion with deterministic tie-breaking = identical queries always return identical results.<p>Photo/Video RAG -- Index your photo library with OCR, captions, GPS binning, per-region embeddings. Query &quot;find that receipt from the restaurant&quot; searches text, visual similarity, and location simultaneously. Videos get segmented with keyframe embeddings + transcript mapping. Results include timecodes for jump-to-moment navigation. All offline -- iCloud-only photos get metadata-only indexing.\nSwift 6.2 strict concurrency -- Every orchestrator is an actor. Thread safety proven at compile time, not runtime. Zero data races, zero @unchecked Sendable, zero escape hatches.<p>Deterministic context assembly -- Same query + same data = byte-identical context every time. Three-tier surrogate compression (full/gist/micro) adapts based on memory age. Bundled cl100k_base tokenizer = no network, no nondeterminism.<p>import Wax<p>let brain = try await MemoryOrchestrator(at: URL(fileURLWithPath: &quot;brain.mv2s&quot;))<p>// Index\ntry await brain.remember(&quot;User prefers dark mode, gets headaches from bright screens&quot;)<p>// Retrieve\nlet context = try await brain.recall(query: &quot;user display preferences&quot;)\n// Returns relevant memories with source attribution, ready for LLM context<p>What makes this different:<p>Zero dependencies on cloud infrastructure -- No API keys, no vendor lock-in, no telemetry\nProduction-grade concurrency -- Not &quot;it works in my tests,&quot; but compile-time proven thread safety\nMultimodal from the ground up -- Text, photos, videos indexed with shared semantics\nPerformance that unlocks new UX -- Sub-millisecond latency enables real-time RAG workflows<p>## Wax Performance (Apple Silicon, as of Feb 17, <em>2026</em>)<p><pre><code>  - 0.84ms <em>vector</em> search at 10K docs (Metal, warm cache)\n  - 9.2ms first-query after cold-open for <em>vector</em> search\n  - ~125x faster than CPU (105ms) and ~178x faster than SQLite FTS5 (150ms) in\n    the same 10K benchmark\n  - 17ms cold-open \u2192 first query overall\n  - 10K ingest in 7.756s (~1289 docs/s) with hybrid batched ingest\n  - 0.103s hybrid search on 10K docs\n  - Recall path: 0.101\u20130.103s (smoke/standard workloads)\n</code></pre>\nBuilt for: Developers shipping AI-native apps who want RAG without the infrastructure overhead. Your data stays local, your users stay private, your app stays fast.<p>The storage format and search pipeline are stable. The API surface is early but functional. If you're building RAG into Swift apps, I'd love your feedback.<p>GitHub: <a href=\"https://github.com/christopherkarani/Wax\" rel=\"nofollow\">https://github.com/christopherkarani/Wax</a><p>Star it if you're tired of spinning up <em>vector</em> <em>databases</em> for what should be a library call."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"Sub-Millisecond RAG on Apple Silicon. No Server. No API. One File"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/christopherkarani/Wax"}},"_tags":["comment","author_ckarani","story_47048731"],"author":"ckarani","children":[47051763,47051939,47052322],"comment_text":"I built Wax because every RAG solution required either Pinecone&#x2F;Weaviate in the cloud or ChromaDB&#x2F;Qdrant running locally. I wanted the SQLite of RAG -- import a library, open a file, query. Except for multimodal content at GPU speed.<p>The architecture that makes this work:\nMetal-accelerated vector search -- Embeddings live directly in unified memory (MTLBuffer). Zero CPU-GPU copy overhead. Adaptive SIMD4&#x2F;SIMD8 kernels + GPU-side bitonic sort = sub-millisecond search on 10K+ vectors (vs ~100ms CPU). This isn&#x27;t just &quot;faster&quot; -- it enables interactive search UX that wasn&#x27;t possible before.<p>Atomic single-file storage (.mv2s) -- Everything in one crash-safe binary: embeddings, BM25 index, metadata, compressed payloads. Dual-header writes with generation counters = kill -9 safe. Sync via iCloud, email it, commit to git. The file format is deterministic -- identical input produces byte-identical output.<p>Query-adaptive hybrid fusion -- Four parallel search lanes (BM25, vector, timeline, structured memory). Lightweight classifier detects intent (&quot;when did I...&quot; \u2192 boost timeline, &quot;find documentation about...&quot; \u2192 boost BM25). Reciprocal Rank Fusion with deterministic tie-breaking = identical queries always return identical results.<p>Photo&#x2F;Video RAG -- Index your photo library with OCR, captions, GPS binning, per-region embeddings. Query &quot;find that receipt from the restaurant&quot; searches text, visual similarity, and location simultaneously. Videos get segmented with keyframe embeddings + transcript mapping. Results include timecodes for jump-to-moment navigation. All offline -- iCloud-only photos get metadata-only indexing.\nSwift 6.2 strict concurrency -- Every orchestrator is an actor. Thread safety proven at compile time, not runtime. Zero data races, zero @unchecked Sendable, zero escape hatches.<p>Deterministic context assembly -- Same query + same data = byte-identical context every time. Three-tier surrogate compression (full&#x2F;gist&#x2F;micro) adapts based on memory age. Bundled cl100k_base tokenizer = no network, no nondeterminism.<p>import Wax<p>let brain = try await MemoryOrchestrator(at: URL(fileURLWithPath: &quot;brain.mv2s&quot;))<p>&#x2F;&#x2F; Index\ntry await brain.remember(&quot;User prefers dark mode, gets headaches from bright screens&quot;)<p>&#x2F;&#x2F; Retrieve\nlet context = try await brain.recall(query: &quot;user display preferences&quot;)\n&#x2F;&#x2F; Returns relevant memories with source attribution, ready for LLM context<p>What makes this different:<p>Zero dependencies on cloud infrastructure -- No API keys, no vendor lock-in, no telemetry\nProduction-grade concurrency -- Not &quot;it works in my tests,&quot; but compile-time proven thread safety\nMultimodal from the ground up -- Text, photos, videos indexed with shared semantics\nPerformance that unlocks new UX -- Sub-millisecond latency enables real-time RAG workflows<p>## Wax Performance (Apple Silicon, as of Feb 17, 2026)<p><pre><code>  - 0.84ms vector search at 10K docs (Metal, warm cache)\n  - 9.2ms first-query after cold-open for vector search\n  - ~125x faster than CPU (105ms) and ~178x faster than SQLite FTS5 (150ms) in\n    the same 10K benchmark\n  - 17ms cold-open \u2192 first query overall\n  - 10K ingest in 7.756s (~1289 docs&#x2F;s) with hybrid batched ingest\n  - 0.103s hybrid search on 10K docs\n  - Recall path: 0.101\u20130.103s (smoke&#x2F;standard workloads)\n</code></pre>\nBuilt for: Developers shipping AI-native apps who want RAG without the infrastructure overhead. Your data stays local, your users stay private, your app stays fast.<p>The storage format and search pipeline are stable. The API surface is early but functional. If you&#x27;re building RAG into Swift apps, I&#x27;d love your feedback.<p>GitHub: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;christopherkarani&#x2F;Wax\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;christopherkarani&#x2F;Wax</a><p>Star it if you&#x27;re tired of spinning up vector databases for what should be a library call.","created_at":"2026-02-17T15:43:37Z","created_at_i":1771343017,"objectID":"47048732","parent_id":47048731,"story_id":47048731,"story_title":"Sub-Millisecond RAG on Apple Silicon. No Server. No API. One File","story_url":"https://github.com/christopherkarani/Wax","updated_at":"2026-03-05T23:34:26Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"TeMPOraL"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"Deep Search <i>is</i> RAG - that is, if we're still expanding the acronym instead of treating it as a word that just means &quot;queries a <em>vector</em> <em>database</em>&quot;.<p>Prediction for Next Hot Thing in Q4 2025 / Q1 <em>2026</em>: someone will make the Nobel prize-worthy discovery that you can stuff results of your deep search into a <em>database</em> (<em>vector</em> or otherwise) and then use it to improve the ability to compile a higher-quality report from much larger amount of sources.<p>We'll call it DeepRAG or Retrieval Augmented Deep Research or something.<p>Prediction for Q2 <em>2026</em>: next Nobel prize awarded for realizing you may as well stop treating report generation as the core aspect of &quot;deep research&quot; (as it obviously makes no sense, but hey, time traveler's spoilers, sorry!), and stop at the &quot;stuff search results into a <em>database</em>&quot; and let users &quot;chat with the search results&quot;, making the research an interactive process.<p>We'll call this huge scientific breakthrough &quot;DeepRAG With Human Feedback&quot;, or &quot;DRHF&quot;."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"The Differences Between Deep Research, Deep Research, and Deep Research"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://leehanchung.github.io/blogs/2025/02/26/deep-research/"}},"_tags":["comment","author_TeMPOraL","story_43236184"],"author":"TeMPOraL","children":[43296060,43338107],"comment_text":"Deep Search <i>is</i> RAG - that is, if we&#x27;re still expanding the acronym instead of treating it as a word that just means &quot;queries a vector database&quot;.<p>Prediction for Next Hot Thing in Q4 2025 &#x2F; Q1 2026: someone will make the Nobel prize-worthy discovery that you can stuff results of your deep search into a database (vector or otherwise) and then use it to improve the ability to compile a higher-quality report from much larger amount of sources.<p>We&#x27;ll call it DeepRAG or Retrieval Augmented Deep Research or something.<p>Prediction for Q2 2026: next Nobel prize awarded for realizing you may as well stop treating report generation as the core aspect of &quot;deep research&quot; (as it obviously makes no sense, but hey, time traveler&#x27;s spoilers, sorry!), and stop at the &quot;stuff search results into a database&quot; and let users &quot;chat with the search results&quot;, making the research an interactive process.<p>We&#x27;ll call this huge scientific breakthrough &quot;DeepRAG With Human Feedback&quot;, or &quot;DRHF&quot;.","created_at":"2025-03-06T07:53:26Z","created_at_i":1741247606,"objectID":"43277591","parent_id":43267539,"story_id":43236184,"story_title":"The Differences Between Deep Research, Deep Research, and Deep Research","story_url":"https://leehanchung.github.io/blogs/2025/02/26/deep-research/","updated_at":"2025-04-16T15:43:58Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"serengil"},"story_text":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["vector","database"],"value":"DeepFace previously relied on filesystem-based datasets and pickle caches for face search, which made it stateful and hard to expose via APIs.<p>We introduced register and search functions that store embeddings in a backend DB (Postgres, Mongo, Weaviate, Neo4j).<p>This makes face recognition stateless, horizontally scalable, and suitable for API use.<p>For Postgres and Mongo, FAISS is used for ANN search; Weaviate and Neo4j handle <em>vector</em> indexing natively.<p>Additional <em>databases</em> and <em>vector</em> stores can be added over time."},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["vector"],"value":"Show HN: DeepFace now supports DB-backed <em>vector</em> search for face recognition"},"url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["2026"],"value":"https://sefiks.com/<em>2026</em>/01/01/introducing-brand-new-face-recognition-in-deepface/"}},"_tags":["story","author_serengil","story_46608519","show_hn"],"author":"serengil","created_at":"2026-01-13T21:37:22Z","created_at_i":1768340242,"num_comments":0,"objectID":"46608519","points":2,"story_id":46608519,"story_text":"DeepFace previously relied on filesystem-based datasets and pickle caches for face search, which made it stateful and hard to expose via APIs.<p>We introduced register and search functions that store embeddings in a backend DB (Postgres, Mongo, Weaviate, Neo4j).<p>This makes face recognition stateless, horizontally scalable, and suitable for API use.<p>For Postgres and Mongo, FAISS is used for ANN search; Weaviate and Neo4j handle vector indexing natively.<p>Additional databases and vector stores can be added over time.","title":"Show HN: DeepFace now supports DB-backed vector search for face recognition","updated_at":"2026-03-05T23:23:05Z","url":"https://sefiks.com/2026/01/01/introducing-brand-new-face-recognition-in-deepface/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"fs90"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"Hey HN,<p>I've been working on a multi-model <em>database</em> called NodeDB.<p>Originally, i've found out the idea of SurrealDB quite good. However, it doesn't have some graph and <em>vector</em> features that I need. And since it is just a KV wrapper, instead of purpose-built engine, the performance will never be close to the specialized databases (like Neo4j, Pinecone, Clickhouse, etc).<p>And i've asked myself, what if, there is a <em>database</em> that have the same idea, but built differently? Instead of just treating it as KV <em>database</em>, we build specialized engines for the data.<p>Besides that, I want it to be able to support my IOT/edge project, where i need offline sync capabilities (Currentyl still in progress).<p>Will it work?<p>I put it into test. I've been experimenting and researching for a year, creating multiple versions, and then I created NodeDB.<p>Disclaimer: It is still in public beta (as of May <em>2026</em>), but it really excites me if I can make this db work. And I use AI as assistant for coding and planning. It is nearly impossible to do as a solo developer without any AI assistance.<p>Would love feedback from HN:<p>- Are there specific features or improvements that would make it more useful?<p>If you're interested in experimenting or contributing, the repo is here: GitHub Repo: <a href=\"https://github.com/nodedb-lab/nodedb\" rel=\"nofollow\">https://github.com/nodedb-lab/nodedb</a><p>Looking forward to your thoughts!"},"title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["database"],"value":"Show HN: NodeDB \u2013 High Perfomance Multi-Model <em>Database</em>"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/nodedb-lab/nodedb"}},"_tags":["story","author_fs90","story_48102084","show_hn"],"author":"fs90","children":[48104383,48107469],"created_at":"2026-05-11T23:21:42Z","created_at_i":1778541702,"num_comments":1,"objectID":"48102084","points":5,"story_id":48102084,"story_text":"Hey HN,<p>I&#x27;ve been working on a multi-model database called NodeDB.<p>Originally, i&#x27;ve found out the idea of SurrealDB quite good. However, it doesn&#x27;t have some graph and vector features that I need. And since it is just a KV wrapper, instead of purpose-built engine, the performance will never be close to the specialized databases (like Neo4j, Pinecone, Clickhouse, etc).<p>And i&#x27;ve asked myself, what if, there is a database that have the same idea, but built differently? Instead of just treating it as KV database, we build specialized engines for the data.<p>Besides that, I want it to be able to support my IOT&#x2F;edge project, where i need offline sync capabilities (Currentyl still in progress).<p>Will it work?<p>I put it into test. I&#x27;ve been experimenting and researching for a year, creating multiple versions, and then I created NodeDB.<p>Disclaimer: It is still in public beta (as of May 2026), but it really excites me if I can make this db work. And I use AI as assistant for coding and planning. It is nearly impossible to do as a solo developer without any AI assistance.<p>Would love feedback from HN:<p>- Are there specific features or improvements that would make it more useful?<p>If you&#x27;re interested in experimenting or contributing, the repo is here: GitHub Repo: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;nodedb-lab&#x2F;nodedb\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;nodedb-lab&#x2F;nodedb</a><p>Looking forward to your thoughts!","title":"Show HN: NodeDB \u2013 High Perfomance Multi-Model Database","updated_at":"2026-05-24T15:54:17Z","url":"https://github.com/nodedb-lab/nodedb"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"phillipclapham"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"There is a shortfall to our current approach to agent memory. Right now, we are just collecting flat facts across a flat memory surface and creating vectorized chains of ambiguity, then wondering why when we ask an agent why it did something the best answer we can get is a probabilistic half-hallucinated half-answer that does not address the actual details of the issue, because it is simply pattern matching to find untyped similarities.<p>So I built FlowScript.<p>FlowScript is a typed reasoning graph that your agent builds through tool calls during your everyday work. It is NOT a graph <em>database</em>. What it is is a small set of unique opinionated primitives: things like thoughts, questions, decisions, blockers, and each of those have typed relationships between them. Your agent calls the tools as it works and it builds this typed graph, and then afterwards you can query that structure to get actual deterministic answers using five queries: <i>tensions</i>, <i>blocked</i>, <i>why</i>, <i>whatIf</i>, <i>alternatives</i>.<p>What does this look like in practice? Here's an agent that has been reasoning about <em>database</em> choices for a few sessions:<p><pre><code>  &gt; mem.query.why(&quot;node_postgres_decision&quot;)\n\n  PostgreSQL chosen\n      \u2190 &quot;Need ACID for payment processing&quot;\n      \u2190 &quot;Original requirement: handle refunds atomically&quot;\n      \u2190 &quot;Stripe webhook failures in staging revealed race condition&quot;\n\n  &gt; mem.query.tensions()\n\n  &gt;&lt;[performance vs cost]\n      &quot;Redis: sub-ms reads critical for UX&quot; vs &quot;Redis cluster: $200/mo for 3 nodes&quot;\n</code></pre>\nThe why chain traces back to the original constraint and the tension preserves the actual tradeoff being made. These are things no <em>vector</em> store can do, because they are NOT just flat facts, but are relationships and reasoning chains that are being captured in a deterministic way. Meaning you can actually go back and audit the actual reasoning of your agent, how it evolved over time, and see the actual tensions that were being balanced. No more opaque reasoning that is lost as soon as the polished answer is generated. Try that in any other memory system, I'll wait.<p>Other memory systems, when they come across a tension or a contradiction, for the most part they are just simply deleting that. And that is wrong because that tension is new knowledge. Knowledge that we need to actually keep for auditing and because it tells us about the evolution of the system and its cognition over time. So instead of deleting contradictions, we relate and create named relationships for them. Relationships you can query.<p>Every decision, every tension, every piece of reasoning is being deterministically captured into an audit trail and hash-encoded. Now, not only do you have a deterministic reasoning chain, but that reasoning chain is auditable. You can go back to any point within the time that you have audit logs for and deterministically review and understand the actual reasoning chain that your model was using. Something that no other system can offer. The EU AI Act is going to require exactly this kind of transparency by August <em>2026</em>, and as far as I can tell, FlowScript is the first open source agent memory system that is designed to meet that bar.<p>Try it NOW: Our MCP server in Claude Code or Cursor. Install and check our Get Started guide so you can add one JSON block to your editor config and drop a snippet into your project CLAUDE.md file, then restart. Your AI assistant gets a full set of reasoning tools that actually trace causality.<p><pre><code>  pip install flowscript-agents openai\n</code></pre>\nSee flowscript.org for full setup instructions: &lt;<a href=\"https://flowscript.org/get-started\" rel=\"nofollow\">https://flowscript.org/get-started</a>&gt;<p>Or grab the TypeScript SDK for programmatic use:<p><pre><code>  npm install flowscript-core\n</code></pre>\nThere are drop-in adapters for LangGraph, CrewAI, Google ADK, and more Python agent frameworks. MIT licensed. Open source.<p>Repo: &lt;<a href=\"https://github.com/phillipclapham/flowscript\" rel=\"nofollow\">https://github.com/phillipclapham/flowscript</a>&gt;\nDocs: &lt;<a href=\"https://flowscript.org\" rel=\"nofollow\">https://flowscript.org</a>&gt;\nPython SDK: &lt;<a href=\"https://github.com/phillipclapham/flowscript-agents\" rel=\"nofollow\">https://github.com/phillipclapham/flowscript-agents</a>&gt;"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: FlowScript \u2013 Agent memory where contradictions are features"},"url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/phillipclapham/flowscript"}},"_tags":["story","author_phillipclapham","story_47516792","show_hn"],"author":"phillipclapham","children":[47518627,47538689],"created_at":"2026-03-25T13:01:58Z","created_at_i":1774443718,"num_comments":1,"objectID":"47516792","points":2,"story_id":47516792,"story_text":"There is a shortfall to our current approach to agent memory. Right now, we are just collecting flat facts across a flat memory surface and creating vectorized chains of ambiguity, then wondering why when we ask an agent why it did something the best answer we can get is a probabilistic half-hallucinated half-answer that does not address the actual details of the issue, because it is simply pattern matching to find untyped similarities.<p>So I built FlowScript.<p>FlowScript is a typed reasoning graph that your agent builds through tool calls during your everyday work. It is NOT a graph database. What it is is a small set of unique opinionated primitives: things like thoughts, questions, decisions, blockers, and each of those have typed relationships between them. Your agent calls the tools as it works and it builds this typed graph, and then afterwards you can query that structure to get actual deterministic answers using five queries: <i>tensions</i>, <i>blocked</i>, <i>why</i>, <i>whatIf</i>, <i>alternatives</i>.<p>What does this look like in practice? Here&#x27;s an agent that has been reasoning about database choices for a few sessions:<p><pre><code>  &gt; mem.query.why(&quot;node_postgres_decision&quot;)\n\n  PostgreSQL chosen\n      \u2190 &quot;Need ACID for payment processing&quot;\n      \u2190 &quot;Original requirement: handle refunds atomically&quot;\n      \u2190 &quot;Stripe webhook failures in staging revealed race condition&quot;\n\n  &gt; mem.query.tensions()\n\n  &gt;&lt;[performance vs cost]\n      &quot;Redis: sub-ms reads critical for UX&quot; vs &quot;Redis cluster: $200&#x2F;mo for 3 nodes&quot;\n</code></pre>\nThe why chain traces back to the original constraint and the tension preserves the actual tradeoff being made. These are things no vector store can do, because they are NOT just flat facts, but are relationships and reasoning chains that are being captured in a deterministic way. Meaning you can actually go back and audit the actual reasoning of your agent, how it evolved over time, and see the actual tensions that were being balanced. No more opaque reasoning that is lost as soon as the polished answer is generated. Try that in any other memory system, I&#x27;ll wait.<p>Other memory systems, when they come across a tension or a contradiction, for the most part they are just simply deleting that. And that is wrong because that tension is new knowledge. Knowledge that we need to actually keep for auditing and because it tells us about the evolution of the system and its cognition over time. So instead of deleting contradictions, we relate and create named relationships for them. Relationships you can query.<p>Every decision, every tension, every piece of reasoning is being deterministically captured into an audit trail and hash-encoded. Now, not only do you have a deterministic reasoning chain, but that reasoning chain is auditable. You can go back to any point within the time that you have audit logs for and deterministically review and understand the actual reasoning chain that your model was using. Something that no other system can offer. The EU AI Act is going to require exactly this kind of transparency by August 2026, and as far as I can tell, FlowScript is the first open source agent memory system that is designed to meet that bar.<p>Try it NOW: Our MCP server in Claude Code or Cursor. Install and check our Get Started guide so you can add one JSON block to your editor config and drop a snippet into your project CLAUDE.md file, then restart. Your AI assistant gets a full set of reasoning tools that actually trace causality.<p><pre><code>  pip install flowscript-agents openai\n</code></pre>\nSee flowscript.org for full setup instructions: &lt;<a href=\"https:&#x2F;&#x2F;flowscript.org&#x2F;get-started\" rel=\"nofollow\">https:&#x2F;&#x2F;flowscript.org&#x2F;get-started</a>&gt;<p>Or grab the TypeScript SDK for programmatic use:<p><pre><code>  npm install flowscript-core\n</code></pre>\nThere are drop-in adapters for LangGraph, CrewAI, Google ADK, and more Python agent frameworks. MIT licensed. Open source.<p>Repo: &lt;<a href=\"https:&#x2F;&#x2F;github.com&#x2F;phillipclapham&#x2F;flowscript\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;phillipclapham&#x2F;flowscript</a>&gt;\nDocs: &lt;<a href=\"https:&#x2F;&#x2F;flowscript.org\" rel=\"nofollow\">https:&#x2F;&#x2F;flowscript.org</a>&gt;\nPython SDK: &lt;<a href=\"https:&#x2F;&#x2F;github.com&#x2F;phillipclapham&#x2F;flowscript-agents\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;phillipclapham&#x2F;flowscript-agents</a>&gt;","title":"Show HN: FlowScript \u2013 Agent memory where contradictions are features","updated_at":"2026-03-27T03:06:51Z","url":"https://github.com/phillipclapham/flowscript"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"robeenly"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"I've been working on NeuG (pronounced &quot;new-gee&quot;), an embeddable graph <em>database</em> that follows the same philosophy as sqlite and DuckDB \u2014 in-process, zero configuration, just pip install and query.<p>What's different from existing embedded graph DBs:<p>- Dual-mode: start embedded, flip one line to expose as a network service \u2014 same data, same queries, no migration\n- Built on GraphScope Flex, the engine behind the current LDBC SNB Interactive world record (80k+ QPS)<p>Local benchmark highlights on LDBC SNB SF1 (~3M nodes, 17M edges):<p>Embedded mode vs LadybugDB (Kuzu-based): NeuG wins 8/9 LSQB queries single-threaded vs LadybugDB's best multi-threaded result. 287x on triangle patterns (Q3), 91x on two-hop filtering (Q2).<p>Service mode vs Neo4j: 617 QPS vs Neo4j's 12 QPS on LDBC SNB Interactive \u2014 50.6x throughput. P95 latency 20ms vs Neo4j's 1,728ms.<p>Currently Python only. Node.js bindings and GraphRAG/<em>vector</em> extensions are on the roadmap.<p>Would love feedback \u2014 especially from anyone who's tried K\u00f9zu, LadybugDB, or runs Neo4j in production.<p>GitHub: <a href=\"https://github.com/alibaba/neug\" rel=\"nofollow\">https://github.com/alibaba/neug</a>\nBlog post with full details: <a href=\"https://graphscope.io/blog/tech/2026/04/12/neug-one-engine-two-modes\" rel=\"nofollow\">https://graphscope.io/blog/tech/<em>2026</em>/04/12/neug-one-engine-t...</a>"},"title":{"matchLevel":"none","matchedWords":[],"value":"Show HN: NeuG \u2013 High-performance Embedded graph DB, one line to serve"}},"_tags":["story","author_robeenly","story_47843670","show_hn"],"author":"robeenly","created_at":"2026-04-21T02:03:31Z","created_at_i":1776737011,"num_comments":0,"objectID":"47843670","points":2,"story_id":47843670,"story_text":"I&#x27;ve been working on NeuG (pronounced &quot;new-gee&quot;), an embeddable graph database that follows the same philosophy as sqlite and DuckDB \u2014 in-process, zero configuration, just pip install and query.<p>What&#x27;s different from existing embedded graph DBs:<p>- Dual-mode: start embedded, flip one line to expose as a network service \u2014 same data, same queries, no migration\n- Built on GraphScope Flex, the engine behind the current LDBC SNB Interactive world record (80k+ QPS)<p>Local benchmark highlights on LDBC SNB SF1 (~3M nodes, 17M edges):<p>Embedded mode vs LadybugDB (Kuzu-based): NeuG wins 8&#x2F;9 LSQB queries single-threaded vs LadybugDB&#x27;s best multi-threaded result. 287x on triangle patterns (Q3), 91x on two-hop filtering (Q2).<p>Service mode vs Neo4j: 617 QPS vs Neo4j&#x27;s 12 QPS on LDBC SNB Interactive \u2014 50.6x throughput. P95 latency 20ms vs Neo4j&#x27;s 1,728ms.<p>Currently Python only. Node.js bindings and GraphRAG&#x2F;vector extensions are on the roadmap.<p>Would love feedback \u2014 especially from anyone who&#x27;s tried K\u00f9zu, LadybugDB, or runs Neo4j in production.<p>GitHub: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;alibaba&#x2F;neug\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;alibaba&#x2F;neug</a>\nBlog post with full details: <a href=\"https:&#x2F;&#x2F;graphscope.io&#x2F;blog&#x2F;tech&#x2F;2026&#x2F;04&#x2F;12&#x2F;neug-one-engine-two-modes\" rel=\"nofollow\">https:&#x2F;&#x2F;graphscope.io&#x2F;blog&#x2F;tech&#x2F;2026&#x2F;04&#x2F;12&#x2F;neug-one-engine-t...</a>","title":"Show HN: NeuG \u2013 High-performance Embedded graph DB, one line to serve","updated_at":"2026-04-21T02:44:17Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"MissMajordazure"},"story_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"Alphabet's upcoming policy to force developer verification for all sideloaded Android apps (effective September <em>2026</em>) is being marketed as an anti-malware security update. In reality, it is the death of general-purpose computing on Android and a direct threat to enterprise and military operational security.<p>1. The Compromise of Enterprise and Air-Gapped OpSec\nBy forcing offline/sideloaded apps to pass through Alphabet's Play Protect verification pipeline, highly sensitive enterprise tools and proprietary air-gapped field apps are forced to &quot;ping home&quot; to Mountain View. Alphabet is inserting itself as a mandatory gatekeeper for internal trade secrets. True self-sovereign development environments are dead if local execution requires a cryptographic nod from a trillion-dollar monopoly.<p>2. The Death of Device Ownership\nThis marks the absolute end of the right to compile your own code and run it on hardware you physically own. If a student, a sysadmin, or a privacy advocate cannot execute a local binary without registering their real identity, paying a fee, and handing over signing keys to Google, they do not own the device. They are merely renting execution privileges.<p>3. The Frontline/Geopolitical Catastrophe\nThis is where the policy moves from anticompetitive to actively dangerous. Modern asymmetric warfare, such as the defense of Ukraine, relies heavily on COTS (Commercial Off-The-Shelf) Android devices.<p>Troops rely on custom, rapidly iterated, securely sideloaded APKs for encrypted communications, drone targeting, and artillery mapping (e.g., Kropyva). With this update, deploying new devices on the frontlines will require these highly classified, tactical applications to be subjected to Alphabet's signing APIs. \n- Military OpSec cannot rely on the uptime, <em>database</em>, or policy whims of a foreign tech corporation.\n- Pinging external servers for app authorization in a high-EW (Electronic Warfare) environment is an unacceptable <em>vector</em> for failure and geolocation tracking.<p>Alphabet is effectively introducing a kill switch that dictates who gets to deploy tactical software in active warzones. It is time to stop pretending this is about &quot;user safety&quot; and recognize it as an unprecedented enclosure of global hardware sovereignty."},"title":{"matchLevel":"none","matchedWords":[],"value":"Alphabet's sideloading auth: a kill switch for frontline and enterprise OpSec"}},"_tags":["story","author_MissMajordazure","story_47247588","ask_hn"],"author":"MissMajordazure","created_at":"2026-03-04T14:11:42Z","created_at_i":1772633502,"num_comments":0,"objectID":"47247588","points":2,"story_id":47247588,"story_text":"Alphabet&#x27;s upcoming policy to force developer verification for all sideloaded Android apps (effective September 2026) is being marketed as an anti-malware security update. In reality, it is the death of general-purpose computing on Android and a direct threat to enterprise and military operational security.<p>1. The Compromise of Enterprise and Air-Gapped OpSec\nBy forcing offline&#x2F;sideloaded apps to pass through Alphabet&#x27;s Play Protect verification pipeline, highly sensitive enterprise tools and proprietary air-gapped field apps are forced to &quot;ping home&quot; to Mountain View. Alphabet is inserting itself as a mandatory gatekeeper for internal trade secrets. True self-sovereign development environments are dead if local execution requires a cryptographic nod from a trillion-dollar monopoly.<p>2. The Death of Device Ownership\nThis marks the absolute end of the right to compile your own code and run it on hardware you physically own. If a student, a sysadmin, or a privacy advocate cannot execute a local binary without registering their real identity, paying a fee, and handing over signing keys to Google, they do not own the device. They are merely renting execution privileges.<p>3. The Frontline&#x2F;Geopolitical Catastrophe\nThis is where the policy moves from anticompetitive to actively dangerous. Modern asymmetric warfare, such as the defense of Ukraine, relies heavily on COTS (Commercial Off-The-Shelf) Android devices.<p>Troops rely on custom, rapidly iterated, securely sideloaded APKs for encrypted communications, drone targeting, and artillery mapping (e.g., Kropyva). With this update, deploying new devices on the frontlines will require these highly classified, tactical applications to be subjected to Alphabet&#x27;s signing APIs. \n- Military OpSec cannot rely on the uptime, database, or policy whims of a foreign tech corporation.\n- Pinging external servers for app authorization in a high-EW (Electronic Warfare) environment is an unacceptable vector for failure and geolocation tracking.<p>Alphabet is effectively introducing a kill switch that dictates who gets to deploy tactical software in active warzones. It is time to stop pretending this is about &quot;user safety&quot; and recognize it as an unprecedented enclosure of global hardware sovereignty.","title":"Alphabet's sideloading auth: a kill switch for frontline and enterprise OpSec","updated_at":"2026-03-06T16:45:29Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"quantdiy"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"In my hunt for an excellent open-source ERP, I couldn't help but be disappointed with the available options. No true lightweight, modern solution exists that I felt I should invest my time with. My other open-source project, QuanuX, is about to launch, and I wanted to set it up from the outset with a highly modern, scalable ERP solution.\nI've recently rebuilt the CLI for QuanuX in Go and had a fantastic experience. It quickly occurred to me that I could build my own ERP atop Google's IAM-based role management and Google Workspace\u2014completely eliminating the need for Twilio, WhatsApp, and a million third-party apps via the Workspace API. And since the Workspace API made this all possible, why not build the whole thing on GCP to complete the hat tip?<p>But I had to take it a step further. It's <em>2026</em>\u2014who wants to run a boring ERP when agents or bots can manage it 100x better than humans?<p>Enter SKILL.md and what I call the &quot;Knowledge <em>Vector</em>.&quot; By utilizing Google's GCP data suite (<em>Vector</em>/ScaNN) and Qdrant, GERP turns all SKILL.md files, Unix man pages, and developer docs into a perfect <em>vector</em> graph. This ensures any developer\u2014human or AI\u2014can enter the repository anywhere and instantly understand the laws of physics governing the system via semantic search.<p>Speaking of AI, my favorite tool in GERP is the native Model Context Protocol (MCP) server. It exposes the ERP's CLI tools directly over STDIO, allowing LLMs and AI Agents to natively read the immutable audit logs and trigger complex workflows without expensive context switching.<p>GERP is a headless ERP infrastructure built almost entirely in Go with a GraphQL-ready schema for your frontend integration. To achieve infinite horizontal scale, it uses zero SQL foreign keys. Instead, it uses a &quot;Golden Thread&quot; architecture across 8 perfectly isolated Google Cloud Spanner <em>databases</em>, and is orchestrated by Temporal Sagas to handle distributed ACID compensations.<p>There's a lot to add and a ways to go before it replaces NetSuite, but we aren't that far off!<p>Check it out: <a href=\"https://github.com/quantDIY/GERP\" rel=\"nofollow\">https://github.com/quantDIY/GERP</a>\nThis is MIT Licensed and ready for you to explore.<p>@quantDIY"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["vector"],"value":"GERP \u2013 A Headless ERP Built with Go, Spanner, Temporal, and GCP <em>Vector</em>"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://github.com/quantDIY/GERP"}},"_tags":["comment","author_quantdiy","story_47534245"],"author":"quantdiy","comment_text":"In my hunt for an excellent open-source ERP, I couldn&#x27;t help but be disappointed with the available options. No true lightweight, modern solution exists that I felt I should invest my time with. My other open-source project, QuanuX, is about to launch, and I wanted to set it up from the outset with a highly modern, scalable ERP solution.\nI&#x27;ve recently rebuilt the CLI for QuanuX in Go and had a fantastic experience. It quickly occurred to me that I could build my own ERP atop Google&#x27;s IAM-based role management and Google Workspace\u2014completely eliminating the need for Twilio, WhatsApp, and a million third-party apps via the Workspace API. And since the Workspace API made this all possible, why not build the whole thing on GCP to complete the hat tip?<p>But I had to take it a step further. It&#x27;s 2026\u2014who wants to run a boring ERP when agents or bots can manage it 100x better than humans?<p>Enter SKILL.md and what I call the &quot;Knowledge Vector.&quot; By utilizing Google&#x27;s GCP data suite (Vector&#x2F;ScaNN) and Qdrant, GERP turns all SKILL.md files, Unix man pages, and developer docs into a perfect vector graph. This ensures any developer\u2014human or AI\u2014can enter the repository anywhere and instantly understand the laws of physics governing the system via semantic search.<p>Speaking of AI, my favorite tool in GERP is the native Model Context Protocol (MCP) server. It exposes the ERP&#x27;s CLI tools directly over STDIO, allowing LLMs and AI Agents to natively read the immutable audit logs and trigger complex workflows without expensive context switching.<p>GERP is a headless ERP infrastructure built almost entirely in Go with a GraphQL-ready schema for your frontend integration. To achieve infinite horizontal scale, it uses zero SQL foreign keys. Instead, it uses a &quot;Golden Thread&quot; architecture across 8 perfectly isolated Google Cloud Spanner databases, and is orchestrated by Temporal Sagas to handle distributed ACID compensations.<p>There&#x27;s a lot to add and a ways to go before it replaces NetSuite, but we aren&#x27;t that far off!<p>Check it out: <a href=\"https:&#x2F;&#x2F;github.com&#x2F;quantDIY&#x2F;GERP\" rel=\"nofollow\">https:&#x2F;&#x2F;github.com&#x2F;quantDIY&#x2F;GERP</a>\nThis is MIT Licensed and ready for you to explore.<p>@quantDIY","created_at":"2026-03-26T19:06:53Z","created_at_i":1774552013,"objectID":"47534370","parent_id":47534245,"story_id":47534245,"story_title":"GERP \u2013 A Headless ERP Built with Go, Spanner, Temporal, and GCP Vector","story_url":"https://github.com/quantDIY/GERP","updated_at":"2026-03-26T23:29:51Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"linuxhiker"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"The Call for Papers is from:<p>Thursday, October 9. 2025 until Friday, January 30. <em>2026</em>.<p>Your proposal is expected to be reviewed before:<p>Friday, February 6. <em>2026</em>.<p>We have 6 tracks:<p>Postgres Extensions Day: Dedicated to developers building and maintaining PostgreSQL extensions. Share your work, learn from others, and connect with the extensions community.<p>Dev: Application development with PostgreSQL. Topics include using extensions like pg_<em>vector</em> for enhanced search performance, creating custom data types, building applications, and leveraging PostgreSQL features in your development workflow.<p>Ops: Operational practices for PostgreSQL and related technologies. Covers deployment, monitoring, performance tuning, backup strategies, high availability, replication, and infrastructure management.<p>Essentials: Deep dives into core PostgreSQL functionality. Explore data types, built-in features, query optimization, indexing strategies, and fundamental concepts. If it\u2019s in the documentation or should be, this track explores it thoroughly. Core PostgreSQL knowledge that provides long-term career value.<p>Life and Fun: Nothing about Postgres goes here. Non-technical topics that make us human. Share your hobbies, side projects, and interests outside of <em>databases</em>. Topics include beekeeping, gardening, homesteading, solar projects, IoT experiments, mountain biking, camping, and anything else that brings you joy.<p>Professional Development &amp; Wellness: Career growth and personal well-being. Topics include career advancement, managing burnout, work-life integration, mental health, communication skills, leadership development, and building sustainable careers in technology."},"story_title":{"matchLevel":"none","matchedWords":[],"value":"[dead]"}},"_tags":["comment","author_linuxhiker","story_46722477"],"author":"linuxhiker","comment_text":"The Call for Papers is from:<p>Thursday, October 9. 2025 until Friday, January 30. 2026.<p>Your proposal is expected to be reviewed before:<p>Friday, February 6. 2026.<p>We have 6 tracks:<p>Postgres Extensions Day: Dedicated to developers building and maintaining PostgreSQL extensions. Share your work, learn from others, and connect with the extensions community.<p>Dev: Application development with PostgreSQL. Topics include using extensions like pg_vector for enhanced search performance, creating custom data types, building applications, and leveraging PostgreSQL features in your development workflow.<p>Ops: Operational practices for PostgreSQL and related technologies. Covers deployment, monitoring, performance tuning, backup strategies, high availability, replication, and infrastructure management.<p>Essentials: Deep dives into core PostgreSQL functionality. Explore data types, built-in features, query optimization, indexing strategies, and fundamental concepts. If it\u2019s in the documentation or should be, this track explores it thoroughly. Core PostgreSQL knowledge that provides long-term career value.<p>Life and Fun: Nothing about Postgres goes here. Non-technical topics that make us human. Share your hobbies, side projects, and interests outside of databases. Topics include beekeeping, gardening, homesteading, solar projects, IoT experiments, mountain biking, camping, and anything else that brings you joy.<p>Professional Development &amp; Wellness: Career growth and personal well-being. Topics include career advancement, managing burnout, work-life integration, mental health, communication skills, leadership development, and building sustainable careers in technology.","created_at":"2026-01-22T17:36:45Z","created_at_i":1769103405,"objectID":"46722478","parent_id":46722477,"story_id":46722477,"story_title":"[dead]","updated_at":"2026-03-05T23:23:27Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"gk1"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"What is a <em>Vector</em> <em>Database</em>? (<em>202</em>1)"},"url":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["vector","database"],"value":"https://www.pinecone.io/learn/<em>vector</em>-<em>database</em>/"}},"_tags":["story","author_gk1","story_35826929"],"author":"gk1","children":[35827373,35827433,35827436,35827557,35827628,35827656,35827660,35827679,35827714,35827760,35827782,35827823,35827922,35827995,35828055,35828069,35828072,35828098,35828106,35828445,35828667,35829097,35829580,35829972,35829991,35830004,35830034,35830862,35831384,35831642,35832000,35832016,35833786,35859089],"created_at":"2023-05-05T09:12:28Z","created_at_i":1683277948,"num_comments":202,"objectID":"35826929","points":409,"story_id":35826929,"title":"What is a Vector Database? (2021)","updated_at":"2025-12-14T20:35:08Z","url":"https://www.pinecone.io/learn/vector-database/"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"Fendyfd"},"title":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"Milvus <em>Vector</em> <em>Database</em> in <em>202</em>3"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"https://milvus.io/blog/milvus-in-<em>202</em>3-unprecedented-<em>vector</em>-<em>database</em>-amidst-tech-buzz.md"}},"_tags":["story","author_Fendyfd","story_38875103"],"author":"Fendyfd","children":[38875826],"created_at":"2024-01-05T02:43:54Z","created_at_i":1704422634,"num_comments":0,"objectID":"38875103","points":7,"story_id":38875103,"title":"Milvus Vector Database in 2023","updated_at":"2024-09-20T16:05:51Z","url":"https://milvus.io/blog/milvus-in-2023-unprecedented-vector-database-amidst-tech-buzz.md"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"tim_sw"},"title":{"matchLevel":"none","matchedWords":[],"value":"AI startup Pinecone valued at $700M"},"url":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"https://www.businessinsider.com/chroma-weaviate-pinecone-raise-funding-a16z-index-<em>vector</em>-<em>database</em>-ai-<em>202</em>3-3"}},"_tags":["story","author_tim_sw","story_35352194"],"author":"tim_sw","children":[35352731,35352797,35352837],"created_at":"2023-03-29T03:49:44Z","created_at_i":1680061784,"num_comments":6,"objectID":"35352194","points":13,"story_id":35352194,"title":"AI startup Pinecone valued at $700M","updated_at":"2024-09-20T13:39:05Z","url":"https://www.businessinsider.com/chroma-weaviate-pinecone-raise-funding-a16z-index-vector-database-ai-2023-3"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"codekisser"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"what place do <em>vector</em>-native <em>databases</em> have in <em>202</em>5? I feel using pgvector or redisearch works well and most setups will probably be using postgres or redis anyway."},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["database"],"value":"Show HN: Chroma Cloud \u2013 serverless search <em>database</em> for AI"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://trychroma.com/cloud"}},"_tags":["comment","author_codekisser","story_44944241"],"author":"codekisser","children":[44954200,44954263],"comment_text":"what place do vector-native databases have in 2025? I feel using pgvector or redisearch works well and most setups will probably be using postgres or redis anyway.","created_at":"2025-08-19T17:38:26Z","created_at_i":1755625106,"objectID":"44954123","parent_id":44944241,"story_id":44944241,"story_title":"Show HN: Chroma Cloud \u2013 serverless search database for AI","story_url":"https://trychroma.com/cloud","updated_at":"2026-03-05T22:31:29Z"},{"_highlightResult":{"author":{"matchLevel":"none","matchedWords":[],"value":"blackcat201"},"comment_text":{"fullyHighlighted":false,"matchLevel":"full","matchedWords":["vector","database","2026"],"value":"I have been following the <em>vector</em> <em>database</em> trend back in <em>202</em>0 and I ended up with the conclusion: <em>vector</em> search features are a nice to have features which adds more value on existing <em>database</em> (postgres) or text search services (elasticsearch) than using an entirely new framework full of hidden bugs. You could get way higher speedup when you are using the right embedding models and encoding way than just using the <em>vector</em> <em>database</em> with the best underlying optimization. And the bonus side is that you are using a stack which was battle tested (postgres, elasticsearch) vs new kids (pinecone, milvus ... )"},"story_title":{"fullyHighlighted":false,"matchLevel":"partial","matchedWords":["vector","database"],"value":"Do we really need a specialized <em>vector</em> <em>database</em>?"},"story_url":{"matchLevel":"none","matchedWords":[],"value":"https://modelz.ai/blog/pgvector"}},"_tags":["comment","author_blackcat201","story_37097004"],"author":"blackcat201","comment_text":"I have been following the vector database trend back in 2020 and I ended up with the conclusion: vector search features are a nice to have features which adds more value on existing database (postgres) or text search services (elasticsearch) than using an entirely new framework full of hidden bugs. You could get way higher speedup when you are using the right embedding models and encoding way than just using the vector database with the best underlying optimization. And the bonus side is that you are using a stack which was battle tested (postgres, elasticsearch) vs new kids (pinecone, milvus ... )","created_at":"2023-08-12T11:19:40Z","created_at_i":1691839180,"objectID":"37099031","parent_id":37097004,"story_id":37097004,"story_title":"Do we really need a specialized vector database?","story_url":"https://modelz.ai/blog/pgvector","updated_at":"2024-09-20T14:55:12Z"}],"hitsPerPage":20,"nbHits":167,"nbPages":9,"page":0,"params":"query=vector+database+2026&advancedSyntax=true&analyticsTags=backend","processingTimeMS":17,"processingTimingsMS":{"_request":{"roundTrip":15},"afterFetch":{"format":{"highlighting":2,"total":2},"merge":{"mergeLoop":{"prepareNextHit":2,"total":2},"total":2},"total":2},"fetch":{"query":5,"scanning":8,"total":14},"total":17},"query":"vector database 2026","serverTimeMS":20}
