Engineering Document Search (RAG)
Current role · RAG document search
Search and RAG over engineering and technical documentation — PDF, Word, Excel and scanned files — with hybrid keyword + semantic retrieval behind a FastAPI service.
Ingest
Index
Query
Problem
Engineering documents are spread across PDF, Word, Excel and scanned files — a mix that plain keyword search handles poorly.
Solution
OCR and parsing for every format, hybrid retrieval (sentence-transformers + FAISS), FastAPI pipelines and LLM-assisted answers.
Status
In development at Elektroservis — one search across PDF, Word, Excel and scanned engineering documents.
Overview
Engineering and technical documentation is a mix of PDFs, Word files, Excel sheets and scans. My current work at Elektroservis in Moscow is a search and RAG system that makes all of it searchable in one place.
The ingestion pipeline parses each format and runs scanned documents through OCR before indexing. Retrieval is hybrid: keyword matching catches exact terms, while semantic search — sentence-transformers embeddings in a FAISS index — finds passages that say the same thing in different words. Combining the two is meant to give more relevant answers than either approach alone.
FastAPI pipelines connect parsing, indexing, retrieval and LLM-assisted answers. The system is in active development and the code is private, so the diagram shows the architecture.
What it does
- Indexes PDF, Word and Excel files, including scanned documents via OCR
- Hybrid search: keyword matching plus semantic retrieval with sentence-transformers + FAISS
- FastAPI ingestion and query pipelines
- LLM-assisted answers grounded in the retrieved passages
What's next
Hiring for a full-stack, Python or AI role — or need something like this built? Let's talk.
I'm available immediately — on-site or hybrid in Moscow, or remote; full-time or contract. I reply within 48 hours — faster on Telegram.