DocmateAI
A social platform for healthcare professionals with AI-powered document analysis, real-time messaging, and a Chrome extension for quick access. Built during my time at The Trybe — now deprecated but technically interesting.
What DocmateAI was
A social platform designed for healthcare professionals in Algeria — think LinkedIn but specifically for doctors, nurses, and pharmacists. The core idea: healthcare workers need a space to share knowledge, discuss cases (anonymized), and access medical literature without the noise of general social media.
The twist was AI-powered tools:
- Upload a medical document (research paper, clinical guidelines, drug interaction reference) and the system answers questions about it
- Record a voice note and get a transcript with key points extracted
- Search across all uploaded documents using natural language
This project is now deprecated — the client pivoted away from the social angle — but the technical work was some of the most challenging I've done.
The AI stack
RAG pipeline (FastAPI + pgvector)
The document analysis system uses a Retrieval-Augmented Generation pipeline:
Document upload
↓
Text extraction (PDF → text, DOCX → text)
↓
Chunking (512 tokens, 50 token overlap)
↓
Embedding (OpenAI text-embedding-3-small)
↓
Storage in pgvector (PostgreSQL extension)
↓
User query
↓
Embed query → find top-K chunks → feed to GPT-4 as context
↓
Response with source citations
Chunking strategy: Documents are split into 512-token chunks with 50-token overlap. The overlap ensures that information at chunk boundaries isn't lost. Each chunk stores:
- The text content
- The embedding vector (1536 dimensions)
- Metadata (document ID, page number, section heading)
- A checksum (to detect duplicate chunks)
Query processing: When a user asks a question:
- The question is embedded using the same model
- pgvector performs a cosine similarity search across all chunks
- The top 5 chunks are retrieved
- These chunks are fed to GPT-4 as context with the user's question
- GPT-4 generates a response that cites the source document
The response includes inline citations like [Source: WHO Malaria Guidelines, p.12] so the user can verify the answer.
Performance: At 10,000 chunks (roughly 200 documents), the vector search takes about 50ms. At 100,000 chunks, it's around 200ms. pgvector's HNSW index handles this well.
Speech-to-text (Whisper)
Doctors could record voice notes instead of typing. The pipeline:
- Audio is uploaded (WebM or WAV format)
- FastAPI receives the file and sends it to OpenAI's Whisper API
- Whisper transcribes the audio (supports Arabic, French, and English)
- The transcript is processed through the RAG pipeline if the user wants
- Key points are extracted using GPT-4
The Arabic transcription was surprisingly good. Whisper handles Darija (Algerian Arabic) better than expected — it occasionally struggles with very colloquial phrases but gets medical terminology right, which is what matters.
PII safety
Before any document gets embedded, a preprocessing step strips patient-identifiable information:
- Names (matched against a common Algerian names dictionary + pattern matching)
- National ID numbers (regex for Algerian NIS format)
- Phone numbers
- Dates of birth
- Hospital/clinic names (matched against a facility directory)
This was a hard requirement. Medical professionals won't use a tool that might expose patient data to an AI model. The PII stripping runs synchronously during upload — the user sees a "processing" state while it happens.
The stripped fields are replaced with placeholders ([PATIENT_NAME], [ID_NUMBER]) so the document remains readable. The original text is never sent to OpenAI — only the stripped version is embedded.
Real-time features
Messaging: Laravel Echo + Pusher for real-time messaging between users. The WebSocket layer handles:
- Direct messages between two users
- Group conversations (for departments or special interest groups)
- Typing indicators
- Online status
The Chrome extension connected to the same WebSocket channel so doctors could get notifications while browsing medical journals on PubMed or UpToDate.
Notifications: A separate notification system (not just messaging) that alerts users when:
- Someone replies to their post
- A document they follow gets updated
- A new document is uploaded in their specialty
- They have unread messages
Notifications are delivered via WebSocket in real-time and via email as a daily digest.
The Chrome extension
Built with Manifest V3, the extension:
- Adds a sidebar to PubMed and UpToDate pages
- When a doctor is reading a paper, they can click "Analyze with DocmateAI"
- The extension extracts the paper's title and abstract
- Sends it to the DocmateAI backend
- The backend finds related documents in the user's library and generates a summary
- The summary appears in the sidebar
This was the most popular feature according to usage analytics. Doctors loved being able to get context from their own document library without leaving the page they were on.
What I learned
The hardest part wasn't the AI — it was building trust with healthcare professionals. They're (rightly) skeptical about uploading patient-related documents to a platform. Three things were essential for adoption:
- PII stripping — visible and documented. Users could see exactly what was stripped before upload
- Data residency — all data stayed on our servers in Algeria, never sent to US/EU data centers (except for OpenAI API calls, which were covered by OpenAI's data processing agreement)
- Audit trail — every document access, every AI query, every export was logged. Users could see who accessed what and when
This project taught me that the technical implementation is often the easy part. Getting users to trust the system with sensitive data requires transparency and security practices that actually have substance behind them.
Technical notes
The backend was split into two services:
- Laravel — the main application (user management, social features, messaging, document management)
- FastAPI — the AI service (embedding, RAG queries, speech-to-text, PII stripping)
They communicated via internal HTTP calls (no message queue between them — the AI service was request/response, not async). In production, both services ran on the same DigitalOcean droplet with Nginx proxying /api/ai/* requests to FastAPI.
The database was PostgreSQL with the pgvector extension. I used a single database for both services — Laravel handled the application tables, FastAPI handled the embedding vectors. This avoided data synchronization issues but meant the AI service needed direct database access.
Now deprecated, but the codebase is a reference I still pull from for RAG pipeline patterns.
Vous avez un projet similaire ?
Expliquez-moi où vous en êtes et ce qui vous bloque. Réponse rapide sur WhatsApp, devis gratuit.