GreenifyAI logo

Case study · Real estate

GreenifyAI turned 928 labeled listings into 2,000% more sustainability coverage, for a fraction of the cost.

August 4, 2026

GreenifyAI runs greenify.ai, a real estate platform that grades MLS listings for sustainability, solar, insulation, EV charging, and efficient systems, so buyers can filter for eco-friendly homes. Grading had always meant one LLM call per listing. They had used it on 928 of their 113,398 listings.

The way it worked

One LLM call per listing

Reliable, but every listing means another request. Across 63 active days of logs, their busiest single day covered 267 listings, the most they ever pushed through at once.

The way it works now

One model, trained on labels they already had

The 928 listings the LLM had already graded became training data. EasyDeploy learned the pattern, then scored the rest of the catalogue in a single batch job.

GreenifyAI listing cards on greenify.ai showing Sustainable and Partially Sustainable score badges on real MLS listings

Listing cards on greenify.ai, tagged with the sustainability score the trained model now generates.

01 / Coverage

928 scored listings became 19,484.

Coverage of the sustainability score, measured against GreenifyAI's real 113,398 listing catalogue.

Before

928

listings scored — 0.82% of the catalogue

2,000%

more coverage

After

19,484

listings scored — 17.2% of the catalogue

02 / Speed

Their best day versus one batch job.

The highest number of distinct listings the existing LLM scoring service covered in a single day, taken from its own request logs.

Before

267

listings on their busiest single day

6,850%

more scoring throughput

After

6.9 sec

to score all 18,556 backlog listings

03 / Cost

An LLM bill grows with the catalogue. A trained model does not.

Every listing sent through the LLM adds to the invoice, priced per call, with no ceiling. A model trained on labels you already have runs against a fixed prediction allowance, so the bill barely moves as the catalogue grows.

$12

$0

$230

$0

$1,405

$49

928

18.5K

113K

LLM, per call
EasyDeploy, prediction pack

96.6%

lower cost than standard per-call LLM pricing, across the full 113,398 listing catalogue.

The real 18,556 listing run cost nothing. It landed inside the free tier, 5 lifetime training credits and 100k lifetime predictions. The $49 line is what scoring the entire current catalogue costs: one 500K prediction pack, priced once, not per listing.

Scoring the full catalogue: about $1,405 in per-call LLM pricing versus a single $49 prediction pack — roughly 28 times cheaper.

04 / Reliability

Checked against 162 listings the model never saw.

These 162 listings were set aside before training and never used to tune the model. The numbers below come from scoring them once, at the end, and comparing the results to the grades they already had.

Accuracy

69.8%

Against 42.6% for always guessing the most common grade.

Macro F1

0.705

Balanced across all three grades, including the smallest.

Worst case errors

0

Never once confused most sustainable with least sustainable.

Sustainability grades run 1 (least sustainable) to 3 (most). Every miss landed on the grade right next door, never mistaking a top-rated listing for a bottom-rated one.

05 / Process

Three steps, start to finish.

No feature engineering by hand, no hyperparameter tuning, no ML engineer required.

1

Prepare

808 of the existing scores came with enough listing detail to learn from. 20 percent held back, untouched, to check the model honestly.

646 train / 162 held out

2

Train

EasyDeploy searched model families, selected features, and tuned hyperparameters automatically.

6 min 30 sec

3

Score

One batch job against all 18,556 backlog listings, roughly 69 times their busiest day of manual scoring.

6.9 seconds, measured

Your labels are already sitting in production.

If something in your product already makes a judgment call, whether that is a model, a service, or a team, EasyDeploy can learn it and run it at whatever scale you need. Bring a CSV. Have a working model before your next meeting.

Holdout accuracy measured on 162 listings excluded from training and model selection. The 18,556 listing backlog scoring is a completed batch run: grade distribution and 6.9 second runtime are measured outputs. Coverage compares 928 previously scored listings against 19,484 scored after the run, within the real 113,398 listing catalogue (0.82% to 17.2%). The 267 listing comparison is the highest number of distinct listings the existing LLM scoring service covered in a single day, taken from its own request logs (63 active days recorded in total). LLM costs are priced from observed token counts at standard per-call API rates, about $0.0124 per listing, with no batch discount applied; the full-catalogue figure applies that measured rate to all 113,398 current listings — a real catalogue size, not a growth projection. EasyDeploy figures reflect published pricing: the 18,556 listing run was completed within the free tier (5 lifetime training credits, 100k lifetime predictions), and the $49 figure is a single 500K prediction pack, which covers the full catalogue with headroom.

GreenifyAI is at greenify.ai · admin@greenifyai.com.