Pick a generation
Branding by version
1.0
POPPY 1.0
Foundation
Poppy 1.0
—
Identity kit
Brand
Voice & hero
Message
What else should occur this version
Beyond the brand
Intelligence strategy
Other AI for Poppy — foundation & execution models
Poppy shouldn't be locked to one model forever. The principle: best model per task, brand-neutral, measured. Claude is today's brain; we evaluate others as foundation (reasoning) and execution (specialized) options.
| Model family | Best-fit role for Poppy | Use as |
|---|---|---|
| Anthropic Claude | Primary reasoning, tool-use, long context, safety — current brain | Foundation (default) |
| OpenAI GPT | Strong general reasoning + function calling; good fallback / A-B | Foundation (alt) |
| Google Gemini | Very long context + native multimodal (docs, images, video) | Foundation (alt) |
| Llama / Mistral (open) | Cheap, fast, self-hostable for routing, drafts, high-volume tasks | Execution (cost tier) |
| Embeddings (Voyage / OpenAI) | Semantic search/retrieval over 282 tables + docs | Execution (retrieval) |
| Speech (TTS/STT) | Voice in/out for 3.0 — Whisper-class STT + neural TTS | Execution (voice) |
| Vision / OCR | Listing photos, utility bills, doc extraction (utility-extract today) | Execution (vision) |
How we compare LLMs over time
An always-on eval harness
- Golden task sets per job type (ticket draft, listing Q&A, term lookup, summarize) with known-good answers
- Shadow + A/B runs: route a % of real tasks to a challenger model in the background
- Score each on quality · cost · latency · safety; log to a new poppy_model_evals table
- Leaderboard per task type; the router reads the winner and self-tunes within spend caps
- Human spot-check + thumbs feedback feeds the score; never auto-switch on irreversible tasks
The router
poppy_model_routes
- Map task_type primary + fallback model, modality-aware
- Vision tasks multimodal; long docs long-context; bulk cheap tier
- Real metering against spend_caps; hard-stop + alert at 90%
- Brand-neutral: pick on merit, not vendor loyalty
- Every routing decision logged + reversible
Make it better
More ways to improve the roadmap
Trust & proof
Earn belief
- Public trust page: gated actions, audit, Fair Housing
- "Why Poppy said this" — show sources every time
- SOC-style security posture as a selling point
Adoption
Pull, not push
- Per-hat onboarding "first 3 things to ask Poppy"
- Win stories: time saved, deals moved
- Referral / ambassador program for power agents
Measurement
Prove value
- North-star metric: tasks completed for users / week
- Adoption, retention, satisfaction (thumbs) dashboards
- Cost-per-successful-task trend down over time
Personality
Memorable, tasteful
- A signature greeting + light real-estate wit (bounded)
- Seasonal/market moments (never partisan)
- Consistent lockup + the generation color
Ecosystem
Reach
- Connectors: FUB, Zoho, Matterport, Maps
- Embed Poppy in partner + brokerage sites
- API so other PURE tools call Poppy
Governance
Stay safe at scale
- Model evals + red-teaming each major
- Bias/Fair-Housing audits on outputs
- Every release stamped + rollback-ready
Canon: poppy.model_strategy + poppy.brand_roadmap. Companions: Brand Roadmap · 2.0 Guide · Docs Index.