Extend an Existing Evidence-Based Organization Profile & Capability Extraction System

via Freelancer ·

Budget / Salary$10–30
TypeFreelance project
LocationRemote
Posted1 hour ago
We are looking for an experienced Python / document intelligence / data modeling engineer to extend an existing evidence-based organization profiling system.

This is not a greenfield project.

We already have a working Python codebase with:

an existing typed Ngo / organization model;
existing organization profile fields;
evidence/provenance concepts;
an evidence-backed capability projection layer;
automated tests;
anti-leakage safeguards;
an existing KRS parser;
an existing canonical organization schema;
downstream decision logic already consuming organization data;
one real structured organization capability already working end-to-end.

Your task is to extend the existing implementation, not replace or redesign it.

The main goal is to make the current organization profile useful enough for real decision-making by safely adding the most important missing organization capabilities from structured sources, statutes, questionnaires and other supplied documents.

What already exists

The current organization model already contains fields such as:

organization type;
registration date / year;
statutory purposes;
activity domains;
target groups;
geographic capacity;
institution access;
organization experience;
staff capabilities;
staff experience;
program assets.

The existing system also already includes:

provenance-aware candidate concepts;
UNKNOWN semantics;
fail-closed behavior;
deterministic tests;
safeguards preventing unrelated free text from silently becoming a stable organization fact;
downstream integration with an existing decision engine.

You will receive the relevant source files, models, interfaces, tests and examples.

You are not expected to design the organization model from scratch.

Main Task

Your task is to extend the current evidence-backed organization capability layer so that it can safely populate a bounded set of high-value organization profile fields.

Possible priority fields include:

statutory purposes;
activity areas;
target groups;
geographic operating capacity;
organization experience;
institution / facility access;
staff capabilities;
staff experience or qualifications;
program or operational assets;
registration-related facts where authoritative evidence is available.

The final v0.1 scope will be agreed before implementation.

We do not expect support for every possible organization fact.

Input Sources

The module will work with an existing versioned intake contract:

ORGANIZATION_INTAKE_V1

Inputs may include:

organization name;
Polish KRS registration number;
structured questionnaire answers;
statute / governing document;
optional additional documents;
structured facts already available in the existing system.

The exact schema and existing code interfaces will be provided.

Evidence Status

One of the most important requirements is to keep different evidence levels clearly separated.

The system should distinguish between:

VERIFIED — supported by an authoritative structured source;
DOCUMENT_EVIDENCED — explicitly supported by an organization document;
DECLARED — supplied through a structured questionnaire;
UNKNOWN — insufficient evidence.

For example:

Structured registry field
→ VERIFIED

Statute, page 5
→ DOCUMENT_EVIDENCED

Questionnaire answer
→ DECLARED

No reliable evidence
→ UNKNOWN

A declaration must never silently become VERIFIED.

Evidence-First Processing

Conceptually:

Structured source / statute / questionnaire / supplied document
→ evidence-backed candidate
→ validation
→ existing organization model

Every positive fact should preserve appropriate provenance, such as:

source type;
document identifier/hash;
page number where applicable;
supporting evidence quote;
source field/value;
extractor or projector version;
evidence status.

Provenance should remain separate from stable organization semantics where appropriate.

Important Anti-Leakage Rule

The system must not derive stable organization capabilities from unrestricted unrelated free text.

For example, information found only in:

previous application text;
project descriptions;
marketing copy;
notes;
historical free-text profiles;
unrelated documents;

must not silently become a VERIFIED or stable capability.

If a fact cannot be safely established, the correct result is:

UNKNOWN

not an inferred best guess.

Document Processing

Some supplied documents, particularly statutes and registration-related documents, will be in Polish.

Native Polish fluency is helpful but not mandatory if you are comfortable working with Polish-language documents using modern LLMs, translation tools and provided acceptance cases.

We will provide representative documents and expected outputs.

LLM Usage

We are open to LLM-assisted extraction where it is useful, especially for long statutes or complex documents.

However:

LLM output must be treated as a candidate, not automatically as truth;
structured/schema-constrained output is preferred;
evidence must be preserved;
deterministic validation should be applied where possible;
unsupported or ambiguous information must remain UNKNOWN;
the implementation should not unnecessarily depend on a single LLM provider.

A hybrid approach using deterministic parsing + LLM extraction + validation is welcome.

Important Architectural Constraints

Please do not redesign the existing system.

In particular:

do not replace the existing organization model;
do not create a parallel profile architecture;
do not redesign the database unless explicitly approved;
do not build a frontend;
do not build the intake form;
do not implement n8n orchestration;
do not modify the opportunity/document extraction module;
do not rewrite the downstream decision engine;
do not perform unrelated refactors;
do not promote unsupported facts into stable organization capabilities.

If an existing source does not provide a trustworthy binding for a specific field, report the limitation rather than inventing one.

Expected Deliverables

We expect:

extension of the existing Python capability/profile layer;
implementation of the agreed additional organization capabilities;
integration with the existing Ngo model and existing interfaces;
preservation of VERIFIED / DOCUMENT_EVIDENCED / DECLARED / UNKNOWN semantics;
provenance for positive facts;
deterministic automated tests;
anti-leakage tests;
representative real-document tests;
fail-closed handling of ambiguous or missing information;
concise technical documentation;
a short coverage report describing:
supported fields;
source used for each field;
evidence status;
unsupported cases;
known limitations.
Definition of Done

The project will be considered complete when:

the agreed additional organization profile fields are supported;
existing functionality continues to work;
existing organization type capability behavior is preserved;
every positive fact has traceable evidence/provenance;
VERIFIED, DOCUMENT_EVIDENCED, DECLARED and UNKNOWN information remain clearly distinguishable;
unrelated free-text content cannot silently create stable capabilities;
ambiguous or unsupported information remains UNKNOWN;
automated tests pass;
representative real inputs produce the agreed expected outputs;
the implementation is delivered as a bounded extension to the existing codebase and is ready for independent review.
Out of Scope

This project does not include:

building the organization intake form;
frontend development;
n8n orchestration;
opportunity/grant document extraction;
matching logic redesign;
database redesign;
building a new organization profile system from scratch;
rebuilding the existing platform.
Ideal Candidate

We are particularly interested in developers with experience in:

Python;
document intelligence;
structured information extraction;
NLP;
LLM structured outputs;
schema-driven data modeling;
evidence/provenance systems;
PDF processing;
automated testing / pytest;
working safely inside an existing production-oriented codebase.
When Applying

Please briefly explain:

how you would extend an existing evidence-backed organization profile system without redesigning it;
how you would preserve VERIFIED vs DOCUMENT_EVIDENCED vs DECLARED vs UNKNOWN semantics;
where you would use deterministic extraction versus LLM-assisted extraction;
how you would prevent unrelated free-text data from becoming stable facts;
examples of similar Python/document intelligence work;
your estimated delivery time;
your estimated number of hours for a bounded v0.1 extension.

We are looking for someone who can carefully extend an existing, test-driven evidence-based system, not someone proposing a complete rewrite.
python software architecture nlp large language models (llms)
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.