In Commercial Real Estate, a Plausible Answer From AI Is Not Enough

reprints


Artificial intelligence can now produce work that looks remarkably finished. It can summarize a lease, review a rent roll, draft an investment committee memo and explain a change in occupancy in polished language.

That creates a new problem for commercial real estate firms: A document can look ready long before it is safe to use.

SEE ALSO: ApartmentIQ Secures $25M Follow-On Funding Round
Arunabh Dastidar.
Arunabh Dastidar. Photo: Courtesy Leni.

For an asset manager, controller or acquisitions professional, the test is not whether an answer reads well but whether the person receiving it can defend the work. Where did the number come from? Was it pulled from the current rent roll or an older version? Does “occupancy” mean leased, physically occupied or economically occupied? Were the calculations independently checked? What happens when the system is unsure?

Those questions are becoming more important as AI moves beyond drafting and begins completing longer sequences of work.

Consider a routine request: “Send the Monday operating report.”

To a person, that instruction carries a great deal of unstated knowledge. The report must use the correct reporting period, property list and accounting data. It may need to flag certain variances but ignore others. Some information may be available to the asset management team, but restricted from other recipients. The finished document may also require review before it leaves the company.

The request sounds like one task. In practice, it is a series of decisions.

That distinction matters because each individual step can appear reasonable while the final result is wrong. An AI system may retrieve a document, summarize it, add figures to a spreadsheet and prepare an email. But, if it selects an outdated file or sends the report to the wrong distribution list, the fact that every individual step worked correctly offers little comfort.

Commercial real estate is particularly unforgiving of this kind of error. A report can be professionally written and still contain a bad entity mapping, an incorrect date or a formula placed in the wrong cell. A lease summary may accurately quote the document while overlooking the provision that matters most to the owner. An underwriting memo can appear coherent even though it started with an incorrect assumption.

That is why the industry needs to distinguish between generating an answer and producing accountable work. The latter requires more than a strong AI model but a record of which sources were used, which calculations were checked, what changed from the prior period, where uncertainty remains and which decisions still belong to a person.

In other words, trust is not a confidence score placed beside an answer. A system stating that it is “92 percent confident” does not tell an investment committee whether the underlying net operating income was reconciled to the accounting system. It does not tell an asset manager whether the latest lease amendment was included. It does not tell a controller who approved an exception.

Trust comes from evidence. That may include checking calculations against the source workbook, confirming that records are current, identifying discrepancies and stopping when information conflicts. For consequential actions, it also means preserving human approval rather than allowing the system to decide for itself when a report should be distributed, a record changed or an exception accepted.

The way to achieve that is not to ask one AI model to handle every part of the job. Instead, the work can be divided into smaller, more clearly defined steps. One system may identify the correct source documents. Another may compare spreadsheet cells or recalculate a figure. Another may test whether the request began with a faulty assumption. A separate check can compare the finished result with the accounting system, an approved template or another source of record. When the evidence conflicts or a judgment call remains, the work is sent to a person.

Specialized systems are not necessarily more intelligent than the largest general-purpose models but they can perform a narrower job that can be tested more consistently. A tool designed to compare spreadsheet calculations, for example, can be evaluated against a clear right or wrong answer. The same is true of checking whether a document is current, whether a user has permission to see it or whether a required approval has been obtained.

Our research shows how much these surrounding checks can affect performance. In one benchmark, changing how information was retrieved and connected increased retrieval precision from roughly 72 percent to 92 percent. The incidence of retrieving stale or unauthorized information fell from about 41 out of every 1,000 retrievals to fewer than one. 

The underlying AI model did not change, but the gains came from providing better context, directing each step to the appropriate tool and independently checking the result.

These results illustrate a larger point: Better work does not necessarily come from buying a larger model. It can come from breaking the work into defined steps, using the right tool for each one and checking the final result against something a firm already trusts.

Arunabh Dastidar is the co-founder and CEO of real estate investment platform Leni.