All projects
Case study · 03

AI · Data · Multimodal search

SmartSearch / MAYbe Here

SmartSearch is a multimodal contextual search engine, an academic project done in pairs, designed to ingest and query heterogeneous data — images, documents, structured data — from a text or visual query.

Status
Academic project, done in pairs
My role
Paired work on the architecture, multimodal ingestion, vector indexing, search strategy, and load testing.
Stack
Python · LanceDB · Apache Arrow · CLIP · Mistral · Ollama · FastAPI · Streamlit · Docker
03

Searching across 20+ GB of heterogeneous data

01

Context

Finding relevant information across images, tables, PDFs, and text takes more than a keyword index. The project, initially conceived as the technical foundation for a mobile experience called MAYbe Here, gradually refocused on the search engine itself — a pipeline able to cross-reference several forms of information from a text or visual query.

MAYbe Here logo: the stylized initials “Mb?H” with a question mark and a location pin, followed by the full name “MAYbe Here”.
MAYbe Here’s original identity, before the refocus on the SmartSearch engine.

02

Multimodal architecture

The pipeline is organized into independent modules, from source ingestion through to result delivery.

FastAPIStreamlit
  • Ingestion and normalization of heterogeneous sources
  • Persistent storage in a Vector Lakehouse
  • Visual and text embeddings
  • Local semantic arbitration
  • Multi-pass hybrid search and confidence scoring
  • Search API and demo interface

03

Vector Lakehouse

Storage relies on LanceDB and Apache Arrow to unify vectors, metadata, and structured data in a single persistent store. This replaces an initial approach based on in-memory vector indexes, which was too RAM-dependent for large volumes. Images and text are projected into a shared vector space with CLIP, enabling search by visual, textual, or combined similarity.

LanceDBApache ArrowCLIP
SmartSearch Explorer interface showing the multimodal catalogue indexed in LanceDB, with over 3.3 million vectors and a SQL filter.
Vector Lakehouse multimodal catalogue: vectors, metadata, and confidence scores, all queryable in SQL.

04

Multimodal ingestion

The engine accepts several source types — image folders, CSVs and structured data, PDFs, and plain text. Hierarchical scanning detects changes and limits re-indexing to items that actually changed; processing is done in streaming batches to keep memory usage under control on large corpora.

  • Change detection and incremental re-indexing
  • Streaming, batch-based processing
  • Trust contract and assigned domain per ingested folder
Terminal showing the ingestion pipeline running in a Docker container: environment check, detection of LanceDB and Ollama, dataset scanning, and LLM arbitration for column mapping.
Interface listing trust contracts per ingested folder, with assigned domain and initial confidence score.
A trust contract is established per ingested folder, with an assigned domain and an initial confidence score.

05

Local AI and restraint

The LLM is not applied systematically to every piece of data. Mistral, run locally via Ollama, steps in as an arbiter only when a source’s structure or semantics cannot be reliably determined by simple rules. For structured data, the engine analyzes a sample to work out a mapping plan, then applies it deterministically to the rest of the file — a separation that reserves generative inference for the steps where it actually adds value.

MistralOllama

06

Multimodal search

A query can be analyzed for its visual similarity, its semantic closeness between image and text, and its estimated domain or intent. Results from these different passes are then merged, deduplicated, and ranked using a confidence score — an approach that handles ambiguous visual queries without relying on a single similarity measure.

Search API response to an image query: ranked results with domain, label, confidence score, and a breakdown of the visual, textual, and intent scores.
Result of an image search, with confidence score and a breakdown of the visual, textual, and intent passes.

07

Robustness

Several mechanisms let the engine run beyond a simple demo notebook.

Docker
  • Memory usage monitoring and adaptive batch processing
  • Automatic pausing when resources become too constrained
  • Write retries on temporary locks
  • Environment validation before launching the pipeline
  • Reproducible containerized deployment

08

Validation

The engine was evaluated with large-scale ingestion scenarios and search tests, on datasets representing more than 20 GB of heterogeneous data and around 100,000 images, complemented by structured and document files. Trials covered both CPU and GPU environments.

  • RAM / VRAM consumption
  • Ingestion time and resuming an already-indexed corpus without full reprocessing
  • Search latency and result relevance
  • Robustness on queries absent from the indexing datasets

09

End-to-end chain

SmartSearch links multimodal ingestion, vector storage, vision, local semantic processing, hybrid search, and result delivery into a single chain. The project focuses less on the final interface than on the technical foundation that turns heterogeneous data into contextualized, traceable results, queryable locally.

Next project

Université de Florence