Perplexity Research (a specialized, high-depth mode within the broader Perplexity AI platform) has released pplx-embed-v2-late, a new family of late-interaction search models engineered to process both text and complex visual documents. Unlike traditional dense retrieval models that compress entire pages into a single point of data, these models retain detailed token-level features. This multi-vector approach allows the system to match specific parts of a query to specific sections of a document, enabling direct search over rendered PDF pages, tables, and images without needing a separate text-extraction pipeline.
In the announcement, they share that the release includes two sizes: a compact lightweight version designed for fast edge deployment and a high-capacity model built for maximum search accuracy. Because both variants are distilled into the exact same representation space, systems can run asymmetric search setups: large document libraries are indexed once using the heavier model for high accuracy, while fast live queries are processed on the fly by the lightweight model.
In their benchmark testing, the release demonstrated strong performance across visual document processing, complex domain-specific text searches, and multi-step AI agent workflows. Both model checkpoints have been published on Hugging Face for integration into standard open-source developer toolkits.