## What’s happening at Agentset.

Stay informed with product updates, company news, and insights on how to sell smarter at your company.

## RAG for the AI SDK

A practical guide to implementing retrieval-augmented generation with the AI SDK — from ingesting documents into vectors to retrieving context at query time.

[Read more](/content/blog/rag-for-the-ai-sdk/index.html)

## Gemini 2 Is the Top Model for Embeddings

Google released Gemini Embedding 2, their first natively multimodal embedding model. We ran it against 17 models across 7 datasets. It takes #1 with 1605 Elo, but the top three are within 18 points of each other.

[Read more](/content/blog/gemini-2-embedding/index.html)

## GPT-5.4: OpenAI's RAG Regression

OpenAI released GPT-5.4 with a Pro variant for extended reasoning. We tested both on our LLM-for-RAG leaderboard across three workloads.

[Read more](/content/blog/gpt-5.4-rag-regression/index.html)

## Zembed-1: The Current Best Embedding Model

ZeroEntropy released zembed-1, a 4B embedding model distilled from zerank-2 reranker. We tested it to see how it performs on real retrieval tasks.

[Read more](/content/blog/zembed-1/index.html)

## Claude Sonnet 4.6 for RAG

We evaluated Claude Sonnet 4.6 in a RAG setup across factual retrieval, synthesis, and scientific tasks compared to frontier models.

[Read more](/content/blog/sonnet-4.6-in-rag/index.html)

## Voyage 4: Evaluation Notes

We tested Gemini 3 inside an actual retrieval setup and compared it directly with GPT-5.1 across five areas that matter for RAG.

[Read more](/content/blog/voyage-4/index.html)

## Claude Opus 4.6 Performance in RAG

We evaluated Claude Opus 4.6 in a RAG setup across factual retrieval, synthesis, and scientific tasks versus 11 frontier models.

[Read more](/content/blog/opus-4.6-in-rag/index.html)

## How to Detect Hallucinations in RAG

RAG helps, but hallucinations still happen. We tested four detection methods to find the best for production—comparing accuracy, latency, and cost.

[Read more](/content/blog/how-to-detect-hallucinations-in-rag/index.html)

## Multimodal vs Text Embeddings: Performance Comparison

We compared a text-based and a multimodal embedding pipeline across text, tables, and charts to see where multimodal actually helps.

[Read more](/content/blog/multimodal-vs-text-embeddings/index.html)

## Gemini 3 Flash: A strong factual RAG model

We evaluated Gemini 3 Flash in a RAG setup to understand where it excels and where it falls short—focusing on factual retrieval and grounding.

[Read more](/content/blog/gemini-3-flash/index.html)

## Cohere Rerank 4: A real upgrade over 3.5

We benchmarked Cohere Rerank 4 Pro and Fast against v3.5 and other rerankers under the same RAG pipeline.

[Read more](/content/blog/cohere-reranker-v4/index.html)

## GPT-5.2 RAG Performance: We Tested It

We plugged GPT-5.2 into our LLM RAG leaderboard and compared it against nine other frontier models under the same RAG pipeline.

[Read more](/content/blog/gpt5.2-on-rag/index.html)

## Best Vector Databases for RAG

We reviewed seven popular vector databases to understand how they differ in deployment, cost, and where they fit in real RAG systems.

[Read more](/content/blog/best-vector-db-for-rag/index.html)

## Opus 4.5 is the new best model for RAG

An evaluation of Opus 4.5 inside a real retrieval setup, compared against Gemini 3 Pro and GPT 5.1 across five behaviors that matter for RAG.

[Read more](/content/blog/opus-4.5-eval/index.html)

## Gemini 3 vs GPT 5.1 for RAG

We tested Gemini 3 inside an actual retrieval setup and compared it directly with GPT-5.1 across five areas that matter for RAG.

[Read more](/content/blog/gemini-3-vs-gpt5.1)

## Embedding models have converged

We compared 13 embedding models across 8 datasets using an LLM judge and ELO scoring. The result: almost all of them perform in the same narrow band.

[Read more](/content/blog/embedding-models-converged/index.html)

## Best Reranker for RAG: We tested the top models

We benchmarked eight leading rerankers to find which performs best for real-world RAG pipelines—comparing speed, accuracy, and relevance.

[Read more](/content/blog/best-reranker/index.html)

## Cohere vs ZeRank: Which Reranker Actually Performs Better?

We compared Cohere v3.5 and ZeRank-1 in a RAG pipeline using a BEIR subset and a custom dataset — analyzing accuracy, latency, and LLM preference.

[Read more](/content/blog/cohere-vs-zerank-comparison/index.html)

## Building Effective RAG Pipelines: A Practical Guide

Learn how to design and implement robust retrieval-augmented generation (RAG) pipelines, from document processing to retrieval optimization.

[Read more](/content/blog/building-effective-rag-pipelines-practical-guide/index.html)

## Is RAG Dead?

OpenAI released the GPT 4.1 models supporting 1M token context window. Gemini supports up to 10M tokens in research. Is the RAG era over?

[Read more](/content/blog/is-rag-dead/index.html)

## Automate Business Workflows with AI Agents

Discover how AI agents can transform business operations by automating complex workflows, reducing manual effort, and improving efficiency.

[Read more](/content/blog/automate-business-workflows-with-ai-agents/index.html)

## Building a Proof-of-Concept RAG System in an Afternoon

A practical guide to quickly building a functional retrieval-augmented generation system to demonstrate the value of AI-powered document search.

[Read more](/content/blog/building-a-proof-of-concept-rag-system-in-an-afternoon/index.html)

## The Art of Document Chunking for LLM Applications

Explore the nuances of effective document chunking strategies for retrieval-augmented generation systems and how they impact LLM performance.

[Read more](/content/blog/the-art-of-document-chunking-for-llm-applications/index.html)

## Parsing PDF Documents at Scale

Learn strategies and techniques to efficiently extract structured information from large volumes of PDF documents for use in AI applications.

[Read more](/content/blog/parsing-pdf-documents-at-scale/index.html)
