> For the complete documentation index, see [llms.txt](https://docs.umbraco.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.umbraco.com/umbraco-cms/develop-with-umbraco/application-code/examine/pdfindex-multisearcher.md).

# PDF Indexes and Multisearchers

If you want to index PDF files and search for them, you will need to use the [UmbracoExamine.Pdf extension package](https://github.com/umbraco/UmbracoExamine.PDF).

## Installation

Install with NuGet:`dotnet add package Umbraco.ExaminePDF`

Installing the package creates a new Examine index called **PDFIndex**, which appears in the **Examine Management** dashboard under the **Settings** section. This index extracts and indexes the text content of PDF files uploaded to the Media section, not their filenames.

![PDFIndex index under Examine Management section](https://1282852327-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FbFfxUcTsPCzuMkpCkZsT%2Fuploads%2Fgit-blob-1160a3a0014e40c95abcdbc04fc96812a1b4f766%2Fpdf-index.png?alt=media)

To get results, search for words that appear inside the PDF. Scanned or image-based PDFs have no extractable text and will return no results.

![PDFIndex Search Results](https://1282852327-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FbFfxUcTsPCzuMkpCkZsT%2Fuploads%2Fgit-blob-63a5161ee525bbe94bd5c7193e527be0fd0b9fcf%2Fpdf-index-search-results.png?alt=media)

## Multi-index searchers

A multi-index searcher is a searcher that can search multiple indexes. The multi-index searcher can be helpful when you want to search both the external and internal indexes.

You can register a multi-index searcher with the ExamineManager on startup:

```csharp
using Examine;
using Umbraco.Cms.Core;
using Umbraco.Cms.Core.Composing;
using Umbraco.Cms.Core.DependencyInjection;
using UmbracoExamine.PDF;

namespace MySite.MyCustomIndex;

[ComposeAfter(typeof(ExaminePdfComposer))]
public class ExamineComposer : IComposer
{
    public void Compose(IUmbracoBuilder builder)
    {
        builder.Services.AddExamineLuceneMultiSearcher("MultiSearcher", new[] {Constants.UmbracoIndexes.ExternalIndexName, PdfIndexConstants.PdfIndexName});
    }
}
```

With this approach, the multi-index searcher will show up in the **Examine Management** dashboard.

![MultiSearcher Search in the "Examine Management" dashboard](https://1282852327-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FbFfxUcTsPCzuMkpCkZsT%2Fuploads%2Fgit-blob-03f9b6d14ab8fddd6bcfdb17f3efe7e2b8d3144c%2Fmulti-searcher.png?alt=media)

The multi-index searcher can be resolved in code from the ExamineManager:

```csharp
if (_examineManager.TryGetSearcher("MultiSearcher", out var searcher))
{
    //TODO: use the `searcher` to search
}
```

{% hint style="warning" %}
The implementation of `IPdfTextExtractor` is `PdfSharpTextExtractor` in this library, which uses `PDFSharp` to extract the bytes to convert to text. The implementation does not deal well with Unicode text, which means when some PDF files are read, the result will be 'junk' strings.

You can replace the `IPdfTextExtractor` using your own composer:

`composition.RegisterUnique<IPdfTextExtractor, MyCustomSharpTextExtractor>();`

The `iTextSharp` library deals with Unicode in a better way but is a paid license. If you wish to use `iTextSharp` or another PDF library, you can swap out the `IPdfTextExtractor` with your own implementation.
{% endhint %}


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://docs.umbraco.com/umbraco-cms/develop-with-umbraco/application-code/examine/pdfindex-multisearcher.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `automate deployments from our CI pipeline` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
