Vectara
Back to blog

Moving beyond text conversion: The future of enterprise search

Enterprise search draws data from documents, charts and diagrams. Vectara is building a successor to Boomerang, our multimodal embedding model, and changing retrieval.

5-minute read timeMoving beyond text conversion: The future of enterprise search

Your company’s most valuable knowledge isn’t some text in a paragraph sitting in a document or in emails. It can be a chart in a quarterly review report, a failure signature in a waveform library, a screenshot buried in a support ticket, or a complex architecture diagram hidden away in presentations.

In times of crisis, when you need a specific failure pattern from a waveform or a trend line from a quarterly review, and your current retrieval system can’t see the visual relationships, that information effectively doesn’t exist. For most enterprises, a text-only retrieval is not reaching the right value.

We are closing this gap and moving past the era where AI just reads. The most effective systems today understand your business through a mix of text, sight, and structure. Therefore, the successor to our multimodal embedding model embeds text and images into one shared space, so the charts and diagrams in your documents are searchable in the same index.

The text-conversion workaround

Until recently, the standard approach was to turn everything into text; OCR the page, caption the chart, summarize the diagram, transcribe the audio, and then index the words. This worked for a while. It made visual data technically searchable. But conversation is compression, with a model in the middle deciding what is worth keeping. When you convert an image to text, you’re compressing it. The problem? If that middleman misses a detail, you’ll never find it.

A caption can note that margins declined in the second half of the year but might fail to mention the exact moment a trend reversed. A text summary of an architecture diagram can list the components but miss the critical connection between the controller and shared memory that your engineer is actually looking for.

This information gap is where enterprise search usually fails. When a user asks a highly specific question, the retrieval system is often looking at a shallow textual interpretation rather than the rich evidence in the source itself. To get real accuracy, the model needs to see the data in its native format. Here’s how multimodal embeddings change the game.

Multimodal embeddings: Search without the middleman

Multimodal models don’t need to convert everything to text. They understand text, images, and charts in the same language.

This means you can type a natural question, and the system compares it directly against the visual evidence. It’s faster, more accurate, and far more comprehensive.

Here is the bottom-line value for your business:

  • Find information across all formats instantly, from scanned PDFs to complex technical drawings.
  • Match charts and diagrams based on the actual concepts they express, not just their filenames.
  • Give your AI agents the full picture, allowing them to cite visual evidence for more reliable answers.

Instead of searching for a description of an image, you’re searching the image itself.

A strategic advantage for the enterprise

For most of their history, embedding models converted text to vectors. Models such as ColPali, VLM2Vec, which retrieve visual documents, existed but they were specialists living outside text retrieval. Training text and images together was assumed to cost in terms of text quality, so combined models were not an option. Today, embedding models from Google, Cohere, Voyage, Amazon, and fast-improving open-weight models embed text and images into a single vector space. The strongest of these models also post the best text retrieval scores on the standard benchmarks at the same time. Enterprises don't care whether the answer is in a paragraph or a chart; they just want the answer. A unified search space allows them to find that answer without knowing which document to click on first. Multimodal embeddings strengthen the retrieval layer by combining multiple representations of the same source. The retrieval system can return the original page or a visual artifact rather than relying on a generated description of it.

Empowering your AI agents

If you’re deploying AI agents for root-cause analysis, financial reporting, or operational support, multimodal retrieval is essential. An agent is only as smart as the context it can access. If relevant context is contained in the charts or diagrams in your repository, a text-only retrieval layer may provide incomplete evidence.

A unified multimodal embedding space allows the retrieval system to search across paragraphs in a manual, diagrams in an architectural document, or scanned pages from an archive without requiring the user to choose a modality in advance.

This creates a more natural and convenient retrieval experience.

Vectara’s next step: Native multimodal retrieval

At Vectara, we believe the retrieval layer should work the way the human brain does by synthesizing text and visuals into a single, cohesive understanding.

We are building a successor to Boomerang, a native multimodal embedding model designed to represent text and images in a shared vector space. It uses a single-vector architecture intended to work with conventional vector-search infrastructure without the vector-volume increase associated with patch-level multi-vector retrieval. We are excited to bring native multimodal embeddings to our platform. It’s a fundamental shift that allows your business to unlock the hidden knowledge through your most complex documents. Your knowledge is multimodal. Your search should be too.

Next, we will look more closely at the strategy and architecture that makes this possible.




Before you go...

Connect with
our Community!