We Didn’t Build a Data Company. We Built an Analytics Company That Needed Better Data
Updated Fri, Sep 4, 2026 - 5 min read
There is a dataset that sits at the heart of every residential property in America. It records what was built, what was changed, what was permitted, and critically, what was not. It spans decades. It covers every jurisdiction in the country. And for most of its existence, nobody knew how to use it.
That dataset is permit data. And it is extraordinarily hard to get right.
The Problem With Permit Data Is the Data Itself

There are a handful of companies that collect building permits at national scale. What most people don’t realize is that collecting the data is the easy part. The hard part is what comes after.
Raw permit data is dirty. Addresses don’t match across sources. Fields are inconsistently populated. Records are duplicated, mislabeled, or missing, not because the permit doesn’t exist, but because the collection failed. A gap in permit history that looks like “no work done” might actually be a jurisdiction that went offline for six months.
This matters because a model built on dirty permit data doesn’t just produce wrong answers; it produces confidently wrong answers. And in underwriting, insurance pricing, and valuation, confidently wrong is worse than unknown.
Most companies that sell permit data sell it as collected. They normalize what they can, fill what they can’t, and ship it. The algorithmic work that makes permit data trustworthy (distinguishing a real absence from a collection gap, catching sub-permit activity that never appears in the public roll, reconciling 5,600+ jurisdictions into a consistent taxonomy) was never done. Because it’s hard, it’s expensive, and there was no forcing function to do it.
Let's connect, and see how we can help you stay ahead of the market.
Contact us
Until you try to build something on top of it.
We Built for Ourselves First
Kukun started as a consumer analytics company. We needed permit data to understand properties, to score their condition, to estimate renovation cost, to build a view of what a home looked like today, not at the time of its last sale.
Once a home is sold, it goes quiet. The MLS snapshot freezes. The appraisal ages. Assessor records lag by years. The only continuous signal of what is actually happening to a home during its ownership is permit activity. We needed that signal to be reliable, or everything we built on top of it would be wrong.
So we built the algorithms. Not to sell data, to use it. We built sub-permit enrichment that surfaces roughly 40% more major residential activity than raw municipal feeds. We built absence intelligence that tells you whether a no-permit result is a risk signal or a coverage gap, and exactly why. We built a normalization layer that makes a permit filed in Los Angeles County and a permit filed in a rural Michigan township speak the same language: 27 consistent project categories, across every source, every jurisdiction.
That took years. It produced PICO™: a 500–850 condition score for 110M+ U.S. properties, the only portable, national, proprietary condition scalar in residential real estate. It produced the only AVM built on permit-enriched records, one that accounts for major renovations that comp-based models never see. It produced ARVE, renovation cost and after-renovation value at zip-code resolution, across 146 project categories and 26.9 billion data points. It produced a contractor network built from permit records, not self-reported profiles.
We ate what we grew. And it worked.
Now We’re Opening the Barn Door
The Kukun Property Knowledge Graph is the harvest made available.
One API call. Any U.S. address. Nine primitives returned, already clean, already joined:
- Permits: 800M+ records, 5,600+ verified sources, all 50 states and territories, median 8-day lag from filing to API
- PICO™ Condition Score: 110M+ U.S. properties, 500–850 scale
- AVM: permit-enriched, renovation-aware, investor-grade
- Renovation Cost & After-Renovation Value (ARVE): 26.9B data points, 146 categories, zip-code resolution
- Permit-Verified Contractor Network: 2.4M+ contractors scored on 22 criteria from permit records
- Systems Remaining Life: per-component useful life for roof, HVAC, and major systems
- Investment Outlook (KIO): zip-code-precision growth forecasts, 1–3 years out
- Market Intelligence: demand and risk signals by zip code and property type
- Property Records: 120M+ U.S. properties, linked to permit history and all scoring models
The join happens inside our infrastructure. The normalization is already done. The absence signals are already scored. What used to take a data team months of preprocessing arrives in a single API response, ready for model ingestion.
For AI-native teams: the Kukun MCP Server makes the full Property Knowledge Graph natively callable by any MCP-compatible agent framework: Claude, ChatGPT, Cursor, LangChain. One configuration step connects any agent to 20+ years of structured property history.
What This Means for Your Team
You don’t have to do the farming. The years of algorithmic work, the source reconciliation, the absence intelligence, the condition scoring- that’s already done.
What you get is what we built for ourselves: property intelligence clean enough to model on, rich enough to see what no comp-based dataset can see, and joined tightly enough that one API call replaces what used to take four vendors and a six-week pipeline.
The harvest is ready. Access it at portal.mykukun.com.
Enterprise licensing and data partnerships: contact@mykukun.com