ComputeLabs Research

Box is integrating Gemini Multimodal Embeddings 2 to index text, images, document pages, spreadsheet tables, and charts together.

· ComputeLabs Research · from the August 18, 2026 edition

Box and Google Cloud are integrating Gemini Multimodal Embeddings 2 into Box’s Agentic Platform. The technology places text, raster images, document pages, rendered spreadsheet tables, and visual charts in a unified multimodal vector space.

The integration is intended to preserve visual and spatial relationships that can be lost when documents are converted into flat text. Examples include maintaining the connection between column headers and data in financial tables, interpreting clinical visuals, and following the logic of multi-page flowcharts.

The system is also intended to support retrieval across mixed enterprise formats, such as a PDF policy, a spreadsheet tracking log, and a presentation deck. This extends text-based retrieval-augmented generation by allowing agents to search visual and structured content alongside narrative text.

Google and Box said enterprises store financial models, clinical-trial protocols, merger-and-acquisition due-diligence materials, engineering schematics, and compliance playbooks in Box. The source describes the architecture and planned capabilities but does not provide an integration-release date, pricing, embedding dimensions, retrieval-accuracy results, or customer deployment figures.

  • Box
  • Gemini Multimodal Embeddings 2

All 20 stories from August 18, 2026