
I started by building an AI skill to help evaluate cars on an auction platform. It eventually became something broader: a decision system combining visual analysis, official vehicle data, market comparables, tax, repair risk, liquidity and capital cost to answer a much more useful question — not whether a car looks interesting, but how much it is rational to pay for it.
I started this project with a practical problem: I was looking at cars on an auction platform and wanted a more reliable way to decide which ones were actually worth buying.
At first, the process seemed straightforward. Open a listing, check the vehicle, review the photographs and damage description, compare the current bid with similar cars on the market, estimate the repair and decide whether to participate. After doing this repeatedly, I realized that finding the information was not the difficult part. The difficult part was connecting it. A low auction price means very little on its own. It has to be considered together with auction fees, VAT, transport, visible damage, possible mechanical risk, registration costs, BPM, market value, expected selling time and the amount of capital that will remain tied up until the car is sold.
The more variables I added, the less useful the original question — “Is this a good car?” — became.
What I actually needed to know was whether a particular vehicle still made economic sense at a particular price, given both what was known and what remained uncertain. More specifically, I wanted to know the highest price I could rationally pay before the economics of the transaction stopped working. That became the basis of the system.
A listing is information, not a decision
Auction platforms are good at presenting what is being sold. They provide specifications, mileage, photographs, descriptions, current bids, pickup locations and sale conditions, but most of the interpretation still happens in the buyer’s head.
A €7,000 bid might be cheap or expensive depending on the commission structure of the auction. A visibly damaged car can still be an attractive opportunity if the repair is predictable, while a cleaner-looking one can be a worse purchase if the mechanical uncertainty is high. Two vehicles with the same expected profit can also be very different uses of capital if one is likely to sell in three weeks and the other may sit in inventory for three months.
I think this distinction matters far beyond vehicle auctions. Products often present information and assume that the user will perform the reasoning, while in many complex decisions the real value lies in understanding the relationships between different pieces of information. I did not need another interface for looking at listings. I needed a repeatable method for turning a listing into a decision.
From a prompt to a system
My first experiments were conversational. I could give a general-purpose AI model a listing and ask it to analyse the vehicle, and the results were often useful.
The problem was consistency. Every new conversation required me to explain what mattered, remind the model which costs to include, challenge assumptions, request market comparisons and then reorganize the answer into something useful enough to compare with another car. A longer prompt could improve the result, but it would not solve the underlying issue. The real problem was not how to phrase the question; it was how to define the method.
The project gradually changed from an AI prompt into a small decision-support system.
Given one or more auction pages, it can discover active passenger-car lots and build a structured record for each vehicle, including its specifications, current bid, listing type, auction description, pickup location and available images. From there, independent analytical layers evaluate the same vehicle from different perspectives. Importantly, those layers do not all use AI. In fact, most of the financial and decision logic is intentionally deterministic.
Using AI where interpretation is useful
The photographs were the most obvious place where AI added value. The system uses a multimodal Vision-Language Model (VLM) to analyse up to three images from an auction listing. Rather than simply detecting objects or damaged areas, the model interprets what it sees in the context of the vehicle and converts visible damage into structured observations: the affected component, location, type of damage, severity and confidence. Those observations then become input for the deterministic part of the system, where they are combined with the auction’s written defect description and translated into repair estimates, uncertainty reserves and the wider risk model.
At the same time, I did not want the model to behave like an imaginary mechanic. A photograph can show that a bumper is damaged and may justify checking what sits behind it, but it cannot prove hidden chassis damage. The system therefore treats directly visible damage differently from suspected structural or mechanical consequences. Structural or critical damage can be flagged for human review, but it is not presented as a diagnosis. Visual confidence is capped, and three auction photographs are explicitly treated as incomplete evidence rather than as a substitute for a physical inspection.The visual findings are then combined with the auction’s written defect description, while duplicated observations are merged rather than counted twice.
The result is not intended to be a workshop quotation. It is a sourcing-oriented repair range combined with an uncertainty reserve. Mechanical risks are modelled separately. An engine problem, gearbox issue, non-starting vehicle or unexplained warning indicator can add its own reserve rather than being hidden inside a general repair number. That distinction became important because visible damage, mechanical risk and missing information represent different kinds of uncertainty and should not be presented as though they were equally well known.
Adding official vehicle data
For cars with a Dutch registration, the system checks official RDW Open Data and adds information such as vehicle identity, first admission, first Dutch registration, APK status, mass, catalogue price, BPM, fuel, CO₂ and Tellerstandoordeel where available. The limitations of the source are part of the design. If a dataset does not provide a complete mileage history, the system does not imply that such a history has been verified.
Vehicles without a Dutch plate require a different path because registration and BPM can materially change the economics. Instead of assigning one generic import reserve, the system can evaluate applicable historical Dutch BPM tariff periods, account for the transition between NEDC and WLTP regimes and apply the relevant depreciation logic.
It can also identify cases where a recognised koerslijst or taxatierapport may be worth investigating. One principle became particularly important here: the system should prefer a supported answer over a more attractive hypothetical one. If a different BPM method appears likely to reduce the tax but the required evidence is not actually available, that potential saving should not silently increase the amount the system is willing to bid.

Building a conservative market view
Valuation turned out to be another area where a seemingly simple task became more complicated once I tried to formalize it. The system searches the Dutch market across AutoScout24, Marktplaats and viaBOVAG while keeping private and dealer inventory separate. Matching begins relatively strictly around model, fuel, body type, transmission, year and mileage, then widens only when the available market is too small.
Damaged cars, export-only offers, lease-price artefacts, duplicates and materially different variants are excluded. Instead of using a broad market average, the system creates two conservative reference points from the cheapest valid matches: Private Market Low and Dealer Market Low. Private Market Low is used as the main basis for expected resale. Dealer pricing is useful as a secondary view of possible upside, but it does not automatically justify paying more for the vehicle. This was a deliberate design decision. The system should be better at identifying reasons not to overpay than at finding optimistic arguments for increasing a bid.

Price is only part of the market
A car can look profitable on paper and still be a poor use of capital if the market for it is slow. For that reason, the wider set of market comparables is also used to estimate liquidity before it is reduced to the final valuation sample.
The system considers factors such as the amount of relevant inventory, source diversity, match quality, price dispersion, the spread between private and dealer asking prices and, where the data exists, the age and freshness of active listings. These signals are combined into a Liquidity Score and an estimated exposure window. The word estimated is intentional. Active listings cannot tell me exactly how long sold vehicles took to find buyers, so this is a proxy rather than historical time-to-sale data. The important part is that this limitation remains visible. The resulting exposure window then becomes part of the financial model because time itself has a cost.
Making time part of the economics
Initially, I thought mainly about purchase price, repair cost and resale value. Once I started thinking in terms of inventory rather than individual transactions, that model became incomplete. A car sitting unsold can create storage, advertising, insurance and tax costs, while the capital invested in it is unavailable for another opportunity. The system therefore converts the estimated exposure period into holding cost and capital cost. It can also apply a liquidity-dependent markdown to the expected resale price rather than assuming that every car will sell immediately at the current market reference. This changes the way two apparently similar opportunities are compared.
A €3,000 expected profit over three weeks is not equivalent to the same €3,000 over three months. That led to an additional measure of Capital Efficiency, which normalizes profit and ROI against time. It is not intended as an investment forecast; it simply makes the cost of waiting visible inside the decision.
Auction fees, tax and the true acquisition cost
Auction fees created another source of potential false precision. There is no useful universal rule such as adding a fixed percentage to every winning bid. Different auctions can have different buyer premiums, VAT bases, handling charges and rules for VAT and Margin vehicles. The system therefore reads the conditions of the specific auction and constructs the relevant fee model from those terms. That calculation also carries confidence. If the auction conditions are ambiguous, the system should not silently guess a percentage and continue to produce a precise-looking MAX BID. In that situation, the calculation is deliberately restricted. VAT is handled in a similar way. Gross cash outlay is kept separate from the net acquisition cost after recoverable purchase VAT, while resale VAT is modelled differently for normal VAT vehicles and Margin vehicles.
The goal is not to recreate accounting software inside the skill. It is simply to prevent a purchasing decision from being based on economics that ignore a material tax effect.
Working backwards to the maximum bid
Eventually these layers meet in one financial model. The system combines a conservative resale value with auction fees, VAT treatment, transport, repair, damage uncertainty, mechanical reserves, BPM or registration costs, expected holding expenses and the cost of capital. From that it calculates Net Trading Profit and Net Trading ROI. Those numbers describe the deal at a given bid, but the more useful output is the Trade MAX BID. Instead of trying to predict where the auction will finish, the system works backwards from the required economics of the finished transaction.
By default, the bid must still allow at least:
€1,500 Net Trading Profit and 15% Net Trading ROI.
The MAX BID is then solved across the complete cost model.
A simplified example makes the idea easier to see. A vehicle might look attractive when the current auction bid is €7,500. Once auction fees, transport, repair risk, BPM, resale tax, expected holding costs and a conservative exit price are included, the model might show that the transaction only meets the target economics up to €8,200. The auction can continue beyond €8,200. My decision should not. That distinction became the most useful output of the entire system. The current bid tells me what other people are currently willing to pay. The MAX BID tells me where my own economics stop working.
Confidence is part of the product
As the model became more sophisticated, confidence became almost as important as the calculated values. A market estimate based on several close comparables should not be presented in the same way as one produced from a widened, low-confidence search. A repair estimate based on three partial photographs should not resemble an inspection report. An unclear auction fee should not become mathematical certainty simply because the output contains two decimal places.
Different parts of the system therefore preserve their own uncertainty and limitations. Some missing information increases a reserve, some reduces a score and some prevents a calculation from being promoted at all. I think this is an important principle for AI products more generally. One of the easiest mistakes is to turn incomplete evidence into polished certainty. For consequential decisions, uncertainty should be designed into the product rather than edited out of it.
From one car to an entire auction
Once the economics of one vehicle were structured, the natural next step was to compare multiple opportunities competing for the same capital. The architecture therefore assigns every sufficiently resolved vehicle a Trade Score from 0 to 100, combining factors such as profit, ROI, dealer upside, condition, capital efficiency, liquidity, registration quality and data completeness. The auction can then be ranked by Trade Score, followed by capital efficiency, current profit, liquidity and remaining headroom to MAX BID. Ranking, however, is still not the same as deciding what to buy.
With limited capital, several smaller opportunities may create a stronger portfolio than one high-scoring but expensive vehicle. The system therefore also includes a portfolio optimizer that can evaluate qualifying cars against a configurable budget. Rather than assuming that each vehicle will be purchased at its current bid, it reserves capital conservatively around the calculated MAX BID and expected holding requirements. A 0/1 knapsack model can then select the combination intended to maximise risk-adjusted expected trading profit within that budget. Architecturally, this means the system is capable of moving from a single vehicle to an auction-level sourcing decision.
In the current ChatGPT prototype, however, performing the full deep-analysis pipeline on every vehicle is not yet practical. That distinction became important later.

When the prototype met the runtime
The system worked, but the environment exposed a different set of constraints. The first was deployment. In the ChatGPT Plus subscription I was using, I could not deploy the skill as the kind of persistent installed capability I initially had in mind. The practical workaround was to create a dedicated ChatGPT Project and upload an archive containing the latest version of the source.
That made the Project itself function as the working environment for the prototype. The bigger issue was execution time.
A complete analysis of a single car currently takes roughly seven to nine minutes, after which the system produces a detailed PDF report for that vehicle. For an experimental analysis of one promising car, that is manageable. For a complete auction, it quickly becomes impractical.
Thirty vehicles processed sequentially could mean several hours of execution, while a larger catalogue would take considerably longer. The architecture might support auction-wide ranking and portfolio optimization, but the current runtime makes brute-force deep analysis of every lot a poor product experience. This was the point where a technical constraint began to influence the product itself.
Depth should be earned
The obvious response would be to make every stage faster, but I think the more interesting answer is architectural. Not every vehicle deserves the same analytical depth. A production system could begin with a fast and relatively inexpensive screening pass that collects the auction, normalizes the listings and eliminates clearly unattractive candidates using information that is already cheap to obtain. Only the most promising vehicles would move into deeper stages such as image analysis, broader market research, BPM scenarios, liquidity modelling and the complete financial calculation. The full PDF report would be produced only for the shortlist.
The flow therefore becomes:
Auction → fast screening → shortlist → deep analysis → MAX BID → final report
rather than running the full pipeline indiscriminately across every vehicle. This is more scalable, but I also think it is a better product principle. Depth should be earned. Expensive analysis is valuable when additional information can realistically change the decision. If a vehicle already fails the basic economics, spending several more minutes proving that conclusion in greater detail creates computational work without creating equivalent user value. Performance constraints therefore became more than an engineering problem. They helped define a better information architecture for the product.
Making the logic testable and versioned
As the system grew, changes in one area began to affect results elsewhere. A change in VAT logic could move the MAX BID. A new liquidity rule could change the expected exit value, which would affect resale tax, profit, ROI, capital efficiency and eventually the vehicle’s position in the auction ranking. At that point, intuition was no longer enough to validate the system.
By version 1.19, the underlying logic was covered by 87 automated tests spanning the major analytical and financial components. Testing was only one part of keeping the project manageable, though. I also learned that explicit versioning needs to start early. When a system contains interconnected assumptions, rules and calculations, even a small adjustment can produce unexpected consequences elsewhere. Without clearly numbered versions, it becomes difficult to understand which implementation produced a particular analysis or to return to a known working state after an unsuccessful experiment.
Keeping explicit versions of the source made it possible to compare behaviour, roll back changes and refer to the exact logic behind a result. That becomes particularly important for a system like this because an output generated by v1.12 and one generated by v1.19 may look visually similar while being based on different VAT logic, scoring rules or assumptions. The version therefore becomes part of the context of the result.
For me, this was another reminder that the intelligence of an AI-enabled product does not live only in the model. Reliability also comes from explicit rules, source provenance, conservative defaults, tests, versioning and the ability to reproduce the reasoning that produced an answer.
What I actually built
I originally thought of the project as an AI skill for analysing auction cars. That description now feels too narrow.
AI is useful, particularly for turning photographs into structured observations, but it represents only one layer of the system. Most of the value comes from combining different forms of evidence into a coherent decision model and defining how the system should behave when that evidence is incomplete.
The harder design questions were about deciding which information belongs in the model, how the variables relate to one another, which assumptions are acceptable, when uncertainty should change the economics and when the system should refuse to produce a stronger conclusion. The output is not simply an answer about a car, it is a structured boundary around a decision.
From AI assistant to decision infrastructure
I think this project reflects a broader direction for AI products. Conversation made general-purpose intelligence easy to access, but many recurring real-world tasks eventually need more structure than a blank input and a good prompt.
They need domain rules, reliable external data, deterministic calculations, explicit treatment of uncertainty and versioned logic that can be tested and reproduced. They also need a clear distinction between the parts where probabilistic AI is genuinely useful and the parts where predictable software is the better tool. Most importantly, they need outputs designed around the decision the user is actually trying to make.
The Auto Auction Catalog does not buy a vehicle for me, and I would not want it to. Its role is to reduce the reasoning that has to be reconstructed for every new listing by gathering the evidence, applying a consistent model, exposing assumptions and identifying the point where the economics stop being rational. The final judgment remains human, but the path to that judgment is much more structured than it was before.

What comes next
The current prototype proved that the analytical model can work, but it also exposed the limits of running this kind of workflow inside a conversational environment. Seven to nine minutes for one deep vehicle analysis is manageable; applying the same process sequentially across an entire auction is not. A dedicated ChatGPT Project works as a useful development environment, but it is not the final product experience I want. I am now working on a web version of the system, where the interface and execution model can be designed around the workflow itself.
That introduces a different set of practical problems: deciding which analysis should happen immediately and which should run asynchronously, how to make a large auction useful before every deep analysis is complete, where expensive results should be cached, how progress and partial results should be shown, what happens when an external data source fails, and how calculations remain reproducible as the underlying logic continues to evolve.
The first version taught me how to structure the decision. The web version is forcing me to answer the next question: how do you turn that reasoning into a product that can operate at scale? That is what I will explore in the next article, including the practical problems I encounter and the solutions I test along the way.


