Senior Python AI Engineer - Private Legal AI Knowledge System (DigitalOcean, OCR, RAG)
Budget / Salary$250–750
TypeFreelance project
LocationRemote
Posted2 hours ago
I need a developer to build Phase 1 of a private AI-powered Legal Knowledge and Case Intelligence System.
This is NOT only a PDF uploader, OCR tool, or chatbot. It must organize, analyze, search, and relate legal documents, communications, evidence, timelines, user notes, and case history into one persistent knowledge system.
EXISTING RESOURCES
I already have DigitalOcean, a custom domain, Google Workspace/Gmail, Google Drive, OpenAI API access, optional Claude/Gemini APIs, OCR/extraction templates, sample legal documents, OCR TXT files, chat/email exports, and extensive project documentation.
DigitalOcean is available and already paid for, but I am NOT requiring a specific hosting provider. You may use another private server/hosting platform for development or production if you can justify it. I must have full administrative access, ownership/control of the deployment, source code, database, files, and credentials.
PHASE 1 GOAL
Phase 1 is for private, non-commercial use and real-world testing. I need the core architecture working reliably, quickly, securely, and cost-effectively.
FILE INGESTION + SOURCE PRESERVATION
Support PDFs/scanned PDFs, images/screenshots, TXT, JSON, DOCX, Excel/CSV, EML/email exports, WhatsApp/Signal/Telegram exports, GPT chat exports, financial/property records, transcripts, and other source files.
Original sources must remain unchanged, indexed, uniquely identified, duplicate-detected, and linked to processed outputs and related matters, people, evidence, communications, events, claims, and notes.
LEGAL OCR / EXTRACTION
Most legal documents are Spanish scanned/image-based PDFs.
PDFs should be rendered into ordered page images before OCR/extraction. Preserve page order, identify blank pages, produce page-level text/metadata, and generate structured TXT/JSON outputs.
Images/screenshots should also be preserved and OCR/extracted.
I already have working extraction templates for Spanish legal documents, metadata, stamps, seals, signatures, handwriting, uncertainty, page order, and source fidelity. Detailed templates will be provided after hiring.
PRIVATE DATABASE + CONNECTED KNOWLEDGE
The system needs its own private database and search layer covering files, documents, cases/matters, people/entities, organizations, courts, attorneys, communications, evidence, events/timelines, claims/assertions, notes/observations, OCR content, ingestion jobs, and errors.
Information must remain connected. Documents, communications, people, events, evidence, claims, notes, dates, institutions, and related matters should link automatically whenever reasonably identifiable.
The goal is for AI retrieval to return complete matter context, not isolated files.
MANUAL INPUT IS FIRST-CLASS DATA
Typed notes, voice-to-text notes, legal strategy, hearing/meeting notes, phone summaries, reminders, observations, and case updates must become timestamped searchable knowledge and link to the appropriate matter/people/events/evidence whenever possible.
PRIVATE AI + SHARED PERSISTENT MEMORY
Provide a private browser-based AI interface.
Persistent memory must come from the shared database/index, not from one temporary chat context.
Every chat should use the same underlying knowledge. If I upload a file, note, email, or case update in one chat, another new or existing chat should be able to retrieve it after ingestion/indexing. Previous chats and user inputs must also be searchable with timestamps.
The AI should retrieve exact source material, answer using stored knowledge, show where information came from, cross-reference records, summarize, identify contradictions/patterns, and locate/download relevant source files.
OpenAI, Claude/Gemini, local models, or a combination may be proposed. Accuracy, speed, source traceability, privacy, cost efficiency, and expandable architecture matter more than a specific model.
Example requests:
“Show me everything related to matter X and give me download links for the source files.”
“Combine the documents and communications for matter X, produce a lawyer brief, and generate a visual timeline.”
“Cross-reference this matter against related matters and identify recurring people, events, arguments, contradictions, or patterns with source references.”
“What did I tell you about this matter last week, and what changed since then?”
FILE MANAGEMENT + CASE PACKAGE EXPORT
Support locating originals, related files, OCR outputs, duplicate detection, missing OCR/unlinked files, case/matter folders, controlled rename/move operations, and downloadable matter packages.
A matter package should be able to include relevant documents, OCR text, communications, evidence, notes, timeline data, manifests, and source references.
EMAIL + COMMUNICATION INTAKE
Support practical Gmail/IMAP/forwarding intake through my custom domain and EML import.
Parse sender, recipients, timestamps, body, attachments, and identifiable relationships to matters/people.
Also ingest WhatsApp/chat exports, preserving participants, message timestamps, media references, key events, evidence indicators, and related entities.
USERS, PRIVACY + AUDIT
Support secure login, admin/regular users, private/shared areas, case-level access, and role-based permissions.
Audit logging is required from Milestone 1 for uploads, OCR, extraction, AI actions, database changes, file actions, imports/exports, errors, and retries.
Failed jobs must not silently disappear. Processing should support status tracking, retries, duplicate-safe/idempotent reprocessing, and clear error records.
STORAGE + BACKUP
Original files must remain manually accessible outside the AI abstraction. Database records must retain exact source/storage references.
Use a practical backup/export strategy. Google Drive is available but should not be a required dependency.
BASIC WEB UI
Simple but usable:
secure login
upload/source browser
search
case/matter views
related records
AI chat
processing status
audit/errors
downloads/exports
TECHNOLOGY
Architecture is flexible. Likely options include Python/FastAPI or Django, PostgreSQL + pgvector or equivalent, background workers/queue, Docker, OpenAI API and optional additional models, and suitable OCR tools/services.
DigitalOcean is available, but you may propose another hosting/server solution. The requirement is a private deployment I fully control and can access/administer.
PHASE 1 SUCCESS
Phase 1 is successful when representative Spanish legal PDFs, images, chat exports, spreadsheets, emails/EML, OCR TXT, and manual notes can be ingested into connected persistent memory; originals remain preserved; AI can retrieve and cross-reference them with source traceability; files can be located/downloaded; and all important processing is auditable.
I should be able to ask:
“Show me everything for matter X.”
“Combine everything for this matter and prepare a lawyer brief with a visual timeline.”
“Cross-reference this matter with related matters and show recurring people, facts, contradictions, or patterns.”
“What did I tell you about this matter last week, and what changed since then?”
KEEP PHASE 1 LEAN
I do NOT need a polished SaaS product, mobile app, advanced multi-agent system, model fine-tuning, or unnecessary enterprise features now.
I need the foundation working well.
CONTINUED WORK
Successful Phase 1 completion will lead to Phase 2 development.
Phase 3 will be the commercial law-firm platform and will include commission/revenue-based compensation tied to secured client contracts and sales under a separate written agreement.
I am looking for a developer who can deliver Phase 1 efficiently and cost-effectively and wants the opportunity to stay involved long-term.
PLEASE ANSWER
Proposed architecture/stack and why.
Realistic Phase 1 timeline and first milestone.
Similar production AI/RAG/OCR systems built.
OCR approach for scanned Spanish legal PDFs.
How originals, OCR outputs, records, and source citations stay linked.
Database/search/vector approach.
How shared persistent memory works across all chats.
Email/chat ingestion approach.
Duplicate/version handling.
Audit, retries, idempotency, security, and deployment approach.
What you need from me after award.
Generic chatbot proposals will not be considered.
This is NOT only a PDF uploader, OCR tool, or chatbot. It must organize, analyze, search, and relate legal documents, communications, evidence, timelines, user notes, and case history into one persistent knowledge system.
EXISTING RESOURCES
I already have DigitalOcean, a custom domain, Google Workspace/Gmail, Google Drive, OpenAI API access, optional Claude/Gemini APIs, OCR/extraction templates, sample legal documents, OCR TXT files, chat/email exports, and extensive project documentation.
DigitalOcean is available and already paid for, but I am NOT requiring a specific hosting provider. You may use another private server/hosting platform for development or production if you can justify it. I must have full administrative access, ownership/control of the deployment, source code, database, files, and credentials.
PHASE 1 GOAL
Phase 1 is for private, non-commercial use and real-world testing. I need the core architecture working reliably, quickly, securely, and cost-effectively.
FILE INGESTION + SOURCE PRESERVATION
Support PDFs/scanned PDFs, images/screenshots, TXT, JSON, DOCX, Excel/CSV, EML/email exports, WhatsApp/Signal/Telegram exports, GPT chat exports, financial/property records, transcripts, and other source files.
Original sources must remain unchanged, indexed, uniquely identified, duplicate-detected, and linked to processed outputs and related matters, people, evidence, communications, events, claims, and notes.
LEGAL OCR / EXTRACTION
Most legal documents are Spanish scanned/image-based PDFs.
PDFs should be rendered into ordered page images before OCR/extraction. Preserve page order, identify blank pages, produce page-level text/metadata, and generate structured TXT/JSON outputs.
Images/screenshots should also be preserved and OCR/extracted.
I already have working extraction templates for Spanish legal documents, metadata, stamps, seals, signatures, handwriting, uncertainty, page order, and source fidelity. Detailed templates will be provided after hiring.
PRIVATE DATABASE + CONNECTED KNOWLEDGE
The system needs its own private database and search layer covering files, documents, cases/matters, people/entities, organizations, courts, attorneys, communications, evidence, events/timelines, claims/assertions, notes/observations, OCR content, ingestion jobs, and errors.
Information must remain connected. Documents, communications, people, events, evidence, claims, notes, dates, institutions, and related matters should link automatically whenever reasonably identifiable.
The goal is for AI retrieval to return complete matter context, not isolated files.
MANUAL INPUT IS FIRST-CLASS DATA
Typed notes, voice-to-text notes, legal strategy, hearing/meeting notes, phone summaries, reminders, observations, and case updates must become timestamped searchable knowledge and link to the appropriate matter/people/events/evidence whenever possible.
PRIVATE AI + SHARED PERSISTENT MEMORY
Provide a private browser-based AI interface.
Persistent memory must come from the shared database/index, not from one temporary chat context.
Every chat should use the same underlying knowledge. If I upload a file, note, email, or case update in one chat, another new or existing chat should be able to retrieve it after ingestion/indexing. Previous chats and user inputs must also be searchable with timestamps.
The AI should retrieve exact source material, answer using stored knowledge, show where information came from, cross-reference records, summarize, identify contradictions/patterns, and locate/download relevant source files.
OpenAI, Claude/Gemini, local models, or a combination may be proposed. Accuracy, speed, source traceability, privacy, cost efficiency, and expandable architecture matter more than a specific model.
Example requests:
“Show me everything related to matter X and give me download links for the source files.”
“Combine the documents and communications for matter X, produce a lawyer brief, and generate a visual timeline.”
“Cross-reference this matter against related matters and identify recurring people, events, arguments, contradictions, or patterns with source references.”
“What did I tell you about this matter last week, and what changed since then?”
FILE MANAGEMENT + CASE PACKAGE EXPORT
Support locating originals, related files, OCR outputs, duplicate detection, missing OCR/unlinked files, case/matter folders, controlled rename/move operations, and downloadable matter packages.
A matter package should be able to include relevant documents, OCR text, communications, evidence, notes, timeline data, manifests, and source references.
EMAIL + COMMUNICATION INTAKE
Support practical Gmail/IMAP/forwarding intake through my custom domain and EML import.
Parse sender, recipients, timestamps, body, attachments, and identifiable relationships to matters/people.
Also ingest WhatsApp/chat exports, preserving participants, message timestamps, media references, key events, evidence indicators, and related entities.
USERS, PRIVACY + AUDIT
Support secure login, admin/regular users, private/shared areas, case-level access, and role-based permissions.
Audit logging is required from Milestone 1 for uploads, OCR, extraction, AI actions, database changes, file actions, imports/exports, errors, and retries.
Failed jobs must not silently disappear. Processing should support status tracking, retries, duplicate-safe/idempotent reprocessing, and clear error records.
STORAGE + BACKUP
Original files must remain manually accessible outside the AI abstraction. Database records must retain exact source/storage references.
Use a practical backup/export strategy. Google Drive is available but should not be a required dependency.
BASIC WEB UI
Simple but usable:
secure login
upload/source browser
search
case/matter views
related records
AI chat
processing status
audit/errors
downloads/exports
TECHNOLOGY
Architecture is flexible. Likely options include Python/FastAPI or Django, PostgreSQL + pgvector or equivalent, background workers/queue, Docker, OpenAI API and optional additional models, and suitable OCR tools/services.
DigitalOcean is available, but you may propose another hosting/server solution. The requirement is a private deployment I fully control and can access/administer.
PHASE 1 SUCCESS
Phase 1 is successful when representative Spanish legal PDFs, images, chat exports, spreadsheets, emails/EML, OCR TXT, and manual notes can be ingested into connected persistent memory; originals remain preserved; AI can retrieve and cross-reference them with source traceability; files can be located/downloaded; and all important processing is auditable.
I should be able to ask:
“Show me everything for matter X.”
“Combine everything for this matter and prepare a lawyer brief with a visual timeline.”
“Cross-reference this matter with related matters and show recurring people, facts, contradictions, or patterns.”
“What did I tell you about this matter last week, and what changed since then?”
KEEP PHASE 1 LEAN
I do NOT need a polished SaaS product, mobile app, advanced multi-agent system, model fine-tuning, or unnecessary enterprise features now.
I need the foundation working well.
CONTINUED WORK
Successful Phase 1 completion will lead to Phase 2 development.
Phase 3 will be the commercial law-firm platform and will include commission/revenue-based compensation tied to secured client contracts and sales under a separate written agreement.
I am looking for a developer who can deliver Phase 1 efficiently and cost-effectively and wants the opportunity to stay involved long-term.
PLEASE ANSWER
Proposed architecture/stack and why.
Realistic Phase 1 timeline and first milestone.
Similar production AI/RAG/OCR systems built.
OCR approach for scanned Spanish legal PDFs.
How originals, OCR outputs, records, and source citations stay linked.
Database/search/vector approach.
How shared persistent memory works across all chats.
Email/chat ingestion approach.
Duplicate/version handling.
Audit, retries, idempotency, security, and deployment approach.
What you need from me after award.
Generic chatbot proposals will not be considered.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.