What’s happening at Agentset.

Stay informed with product updates, company news, and insights on how to sell smarter at your company.

RAG for the AI SDK

A practical guide to implementing retrieval-augmented generation with the AI SDK — from ingesting documents into vectors to retrieving context at query time.

Read more

Gemini 2 Is the Top Model for Embeddings

Google released Gemini Embedding 2, their first natively multimodal embedding model. We ran it against 17 models across 7 datasets. It takes #1 with 1605 Elo, but the top three are within 18 points of each other.

Read more

GPT-5.4: OpenAI's RAG Regression

OpenAI released GPT-5.4 with a Pro variant for extended reasoning. We tested both on our LLM-for-RAG leaderboard across three workloads.

Read more

Zembed-1: The Current Best Embedding Model

ZeroEntropy released zembed-1, a 4B embedding model distilled from zerank-2 reranker. We tested it to see how it performs on real retrieval tasks.

Read more

Claude Sonnet 4.6 for RAG

We evaluated Claude Sonnet 4.6 in a RAG setup across factual retrieval, synthesis, and scientific tasks compared to frontier models.

Read more

Voyage 4: Evaluation Notes

We tested Gemini 3 inside an actual retrieval setup and compared it directly with GPT-5.1 across five areas that matter for RAG.

Read more

Claude Opus 4.6 Performance in RAG

We evaluated Claude Opus 4.6 in a RAG setup across factual retrieval, synthesis, and scientific tasks versus 11 frontier models.

Read more

How to Detect Hallucinations in RAG

RAG helps, but hallucinations still happen. We tested four detection methods to find the best for production—comparing accuracy, latency, and cost.

Read more

Multimodal vs Text Embeddings: Performance Comparison

We compared a text-based and a multimodal embedding pipeline across text, tables, and charts to see where multimodal actually helps.

Read more

Gemini 3 Flash: A strong factual RAG model

We evaluated Gemini 3 Flash in a RAG setup to understand where it excels and where it falls short—focusing on factual retrieval and grounding.

Read more

Cohere Rerank 4: A real upgrade over 3.5

We benchmarked Cohere Rerank 4 Pro and Fast against v3.5 and other rerankers under the same RAG pipeline.

Read more

GPT-5.2 RAG Performance: We Tested It

We plugged GPT-5.2 into our LLM RAG leaderboard and compared it against nine other frontier models under the same RAG pipeline.

Read more

Best Vector Databases for RAG

We reviewed seven popular vector databases to understand how they differ in deployment, cost, and where they fit in real RAG systems.

Read more

Opus 4.5 is the new best model for RAG

An evaluation of Opus 4.5 inside a real retrieval setup, compared against Gemini 3 Pro and GPT 5.1 across five behaviors that matter for RAG.

Read more

Gemini 3 vs GPT 5.1 for RAG

We tested Gemini 3 inside an actual retrieval setup and compared it directly with GPT-5.1 across five areas that matter for RAG.

Read more

Embedding models have converged

We compared 13 embedding models across 8 datasets using an LLM judge and ELO scoring. The result: almost all of them perform in the same narrow band.

Read more

Best Reranker for RAG: We tested the top models

We benchmarked eight leading rerankers to find which performs best for real-world RAG pipelines—comparing speed, accuracy, and relevance.

Read more

Cohere vs ZeRank: Which Reranker Actually Performs Better?

We compared Cohere v3.5 and ZeRank-1 in a RAG pipeline using a BEIR subset and a custom dataset — analyzing accuracy, latency, and LLM preference.

Read more

Building Effective RAG Pipelines: A Practical Guide

Learn how to design and implement robust retrieval-augmented generation (RAG) pipelines, from document processing to retrieval optimization.

Read more

Is RAG Dead?

OpenAI released the GPT 4.1 models supporting 1M token context window. Gemini supports up to 10M tokens in research. Is the RAG era over?

Read more

Automate Business Workflows with AI Agents

Discover how AI agents can transform business operations by automating complex workflows, reducing manual effort, and improving efficiency.

Read more

Building a Proof-of-Concept RAG System in an Afternoon

A practical guide to quickly building a functional retrieval-augmented generation system to demonstrate the value of AI-powered document search.

Read more

The Art of Document Chunking for LLM Applications

Explore the nuances of effective document chunking strategies for retrieval-augmented generation systems and how they impact LLM performance.

Read more

Parsing PDF Documents at Scale

Learn strategies and techniques to efficiently extract structured information from large volumes of PDF documents for use in AI applications.

Read more