The Unstructured Data Problem in Venture Capital
Something like 95 percent of a venture capital firm's knowledge lives in places no system can search: email threads, meeting conversations, board decks sitting in shared drives, and context in someone's head. The structured record captures the other five percent. Every stale pipeline view, every scramble before an LP call, and every repeated question about a portfolio company traces back to this gap.
Where your firm's knowledge actually lives
- Email: deal terms, portfolio updates, introduction requests, LP communications
- Meetings: investment committee discussions, board meetings, founder calls, partner debriefs
- Documents: term sheets, board decks, quarterly reports, fund documents, side letters
- Calendars: who met whom, when, and how often
- Heads: the context that never gets written down at all
- The CRM: contact names, maybe a deal stage, maybe a note from six months ago
The cost of inaccessible knowledge
The cost is not abstract. It shows up as specific, recurring losses:
- Meeting prep: 30 to 45 minutes of manually assembling context from old emails before every meeting.
- Missed signals: a portfolio company's runway dropped to four months in a board deck nobody opened.
- Duplicate work: two partners reach out to the same founder in the same month without knowing.
- Lost history: a partner leaves and three years of relationships and context vanish with them.
- LP questions: "what's the latest on Company X?" triggers a firm-wide scramble instead of a lookup.
- Stale records: the record says Series A; the company raised its B six months ago.
Why CRMs cannot fix this
CRMs cannot fix the unstructured data problem because they are designed for manual input, and venture workflows do not produce structured input naturally. Partners generate knowledge in meetings and email; CRMs need it typed into fields. That mismatch is fundamental, not a training issue.
The consequences are predictable. Adoption follows a decay curve: high in month one, down to roughly 30 to 40 percent of fields maintained by month six. Hiring someone to babysit the record helps but is expensive and does not scale with the volume of communication a firm produces. The work always loses to actual investing, as it should.
What's changing
AI can now read unstructured text and extract structured data reliably, which changes the architecture of the answer. Email parsing, meeting transcription, and document understanding are mature enough to run continuously on a firm's real communication. The shift is from "humans enter data, systems store it" to "AI captures data, humans verify it." That eliminates the adoption problem entirely, because a system that requires no behavior change cannot decay with behavior.
What a solution looks like
A real answer to the unstructured data problem has five properties:
- Connects to existing email and calendar, with no migration project
- Reads and extracts automatically, with no new workflow
- Structures knowledge into connected records: people, companies, deals, updates
- Routes the small fraction of ambiguous extractions to humans for judgment
- Compounds over time, growing institutional memory with every email and meeting
This is what Kosa builds, specifically for venture capital.
Frequently asked questions
What is the unstructured data problem in venture capital?
The overwhelming majority of a VC firm's knowledge, on the order of 95 percent, lives in email threads, meeting conversations, and documents that no system can search or connect. The structured record, usually a CRM, captures the remaining sliver, which is why firm records are chronically stale and incomplete.
Why is CRM adoption so low at VC firms?
Because venture workflows do not naturally produce structured input. Partners communicate in meetings and email, not in database fields, so keeping a CRM current is unpaid extra work. Adoption follows a predictable decay curve: high in month one, down to roughly 30 to 40 percent of fields actually maintained within six months.
What does inaccessible knowledge actually cost a firm?
Concretely: 30 to 45 minutes of manual context assembly per meeting, missed signals like a runway drop buried in an unread board deck, duplicate outreach to the same founder by two partners, scrambles when an LP asks for the latest on a company, and the permanent loss of context when a partner leaves.
How does AI change the unstructured data problem?
AI can now read unstructured text and extract structured data reliably, which inverts the old architecture. Instead of humans entering data for the system to store, the system captures data from communication and humans verify the edge cases. That removes the adoption problem, because no behavior change is required.
Kosa is currently in early access. To see it running on your own pipeline and portfolio, request access at [email protected] or through the form on the homepage.