Field Notes · AI in 2026

The Restrepo AI Field Guide

Best models compared, token economics, local AI vs cloud, open source, control systems, and the privacy stack we deploy. Twenty years in the field, written for owners and operators.

AI in our work · practical, not hype.

AI is the most over-promised feature in our industry. It is also one of the most useful when it is set up right. This is the field guide we wish every client had before they signed with anyone, including us. Models, costs, local versus cloud, open source, control systems, and the privacy stack that keeps a luxury home or a regulated facility clean. Written from twenty years in the field.

Best AI right now · by job, not by hype.

There is no single best. There are tiers, each best at different jobs.

General-purpose chat & reasoning

  • GPT-5 (OpenAI). best all-around. Strong reasoning, tool use, voice. Default for most people.
  • Claude 4.7 Opus / Sonnet (Anthropic). best for long documents, code review, careful writing, proposals.
  • Gemini 3 Pro (Google). best inside Google Workspace, best video understanding, biggest context window.
  • Grok 4 (xAI). best real-time web and X knowledge.
  • Llama 4 / DeepSeek V3 (open weight). best when you need it on-prem or air-gapped. The pick for government and mission-critical facility rooms.

Voice & assistants in the home

  • Apple Intelligence on HomePod / iOS 19. best privacy posture (most processing on-device).
  • Josh.ai. best dedicated luxury home voice. Plays clean with Crestron.
  • Alexa+ (Amazon). best ecosystem breadth, worst privacy posture.
  • Google Assistant w/ Gemini. best for Nest and Google ecosystems.

Image & video generation

  • Sora 2 / GPT-5 image. best photoreal video.
  • Midjourney v7. best art direction and luxury moodboards.
  • Google Veo 3. best long-form video continuity.
  • Runway Gen-4. best for editors who want fine control.
NeedPick
Daily driver chatGPT-5
Writing proposals & contractsClaude 4.7 Opus
Google Docs / Gmail / SheetsGemini 3 Pro
Real-time news, X feedsGrok 4
On-prem / air-gappedLlama 4 or DeepSeek V3
Luxury home voiceJosh.ai or Apple
Hero images for the websiteMidjourney v7
Video for socialSora 2 or Veo 3

Our honest take: most clients use GPT-5 for daily life, Claude for serious writing, and Gemini inside Workspace. That is the stack.

20 new things AI brought · the last 18 months.

  • Real-time voice with conversation latency under 300ms (GPT-5 Voice, Gemini Live).
  • Computer use and agents that drive a browser and finish tasks.
  • On-device models. Apple Intelligence, Pixel Gemini Nano, fully local inference.
  • Million-token context. entire codebases, books, or building manuals in one prompt.
  • Photoreal video. Sora 2, Veo 3, full scenes from text.
  • AI camera analytics that separate person, vehicle, and package and only alert when it matters.
  • AI auto-framing and speaker tracking on Crestron, Q-SYS, Logitech rooms.
  • Live translation in conferencing. Teams, Zoom, Webex, 30+ languages.
  • Auto-generated meeting summaries and action items.
  • Code generation. Claude, Cursor, Copilot Workspace producing real production code.
  • Predictive HVAC and lighting that learns occupancy and sun rhythm.
  • AI Wi-Fi tuning. Ubiquiti and Cisco Meraki self-tune density and roaming.
  • AI-powered IDS on Ubiquiti UISP, Fortinet, Palo Alto.
  • Sentiment analytics tied to room or amenity in hospitality.
  • Voice cloning and dubbing (ElevenLabs, OpenAI). useful and a real fraud surface.
  • Embodied AI. humanoid robotics (Figure, 1X, Optimus) reaching real pilots.
  • AI-driven CAD and BIM. auto room layout, conduit routing, system design assistance.
  • AI search. Perplexity, ChatGPT search, Gemini search replacing parts of Google.
  • Custom GPTs and agents-as-products. every business shipping its own assistant.
  • Hardware AI accelerators in the home (NVIDIA Jetson, Apple Neural Engine, AMD Ryzen AI).

The cost of tokens · where AI bills hide.

A token is roughly 3/4 of a word. Vendors price per million tokens. Two prices for every model: input (what you send) and output (what it writes back). Output is usually 3 to 5x more expensive than input.

ModelInput / 1MOutput / 1MNotes
GPT-5~$3~$15Daily driver.
GPT-5 mini~$0.30~$1.50Cheap and fast.
Claude 4.7 Opus~$15~$75Premium writing and code.
Claude 4.7 Sonnet~$3~$15Strong middle tier.
Gemini 3 Pro~$1.25~$10Cheapest top-tier.
Gemini 3 Flash~$0.10~$0.40Best cost-per-token.
Grok 4~$5~$15Real-time premium.
Llama 4 / DeepSeek (open weight)$0 license$0 licenseYou pay for the GPU.

What a real conversation actually costs

  • Quick chat, 10 turns on GPT-5: ~$0.05
  • Read a 50-page PDF and summarize: ~$0.30
  • Write a 1,500 word blog on Claude Sonnet: ~$0.10
  • One Sora 2 clip, 10 seconds, 1080p: ~$0.50 to $2
  • One Midjourney image: ~$0.04
  • AI voice assistant in a guest room, 50 commands a day: ~$0.20 a day, ~$72 a year per room

Where the hidden costs hide

  • Long context kills you. Sending a million-token codebase 100 times a day is real money.
  • Output dwarfs input. A short prompt asking for a long article costs 5x what you expected.
  • Reasoning modes silently 10x the bill. Thinking tokens are billed even though you never see them.
  • Per-seat AI in business apps. Microsoft 365 Copilot is $30/user/month, Workspace AI is $20. Fifty people = $12K to $18K a year on top of existing licenses.
  • Voice and video. Real-time voice runs ~$0.06 a minute. Video generation $0.50 to $5 a clip. Cheap until you scale.
  • AI camera fees. Cloud VMS like Verkada and Eagle Eye charge $20 to $40 a camera a month. A 30-camera property is $7K to $14K a year.

How to keep AI cost sane

  • Tier the work. Flash and mini for 80%, premium for the 20% that needs it.
  • Cache prompts. Every major vendor supports it now (50 to 90% off repeated context).
  • Local-first where you can. Apple Intelligence and Llama 4 on a Mac mini cost zero per token.
  • Set hard caps on every API key. Every vendor supports it. Most clients get burned because they did not.
  • For personal use, flat-rate seats ($20/month tiers) cover most workflows.
  • For homes, plan ~$200 to $600 a year in AI subscriptions. That is the realistic recurring fee surface today.
  • For commercial, budget AI as OpEx, not a free feature. A real corporate AI bill at a 200 person company is $50K to $200K a year all-in.

Local AI and Ollama · your model on your hardware.

Cloud AI has three permanent problems: your data leaves the building, the bill scales forever, and the internet has bad days. Local AI fixes all three. The model lives on a box you own, in a closet you control, on a network you segment. No tokens. No outage when ChatGPT goes down. No data shipped to a vendor training pipeline.

Ollama in plain language

Ollama is a free, open-source app that runs AI models on your own hardware. Mac, Windows, Linux. Download a model once, run it forever, offline. Type ollama run llama4 and you have a private GPT-5-class model on your kitchen table. What it runs:

  • Llama 4 (Meta). best general-purpose open model
  • DeepSeek V3 / R1. best open reasoning model
  • Mistral / Mixtral. fast, efficient, French
  • Qwen 3 (Alibaba). strong multilingual
  • Gemma 3 (Google). small, fast, embedded use
  • Phi-4 (Microsoft). tiny, runs on a laptop

What it costs

Use caseHardwareOne-timeWhat it runs
Tinker / chatM4 Mac mini 24GB~$1,4007B to 13B snappy
Family workhorseM4 Pro Mac mini 64GB~$2,80070B comfortable
Pro / businessM4 Max Mac Studio 128GB~$4,500Llama 4 70B + DeepSeek R1
Government / seriousRTX 5090 box or DGX Spark~$3K to $10KUp to 200B locally
Edge / embeddedNVIDIA Jetson Orin~$500 to $2KVoice, vision, small LLMs

For comparison, ChatGPT Enterprise is roughly $60/user/month. Ten employees = $7,200/year, every year. A $4,500 Mac Studio pays for itself in about eight months for a 10-person shop and runs free for the next five-plus.

Other ways to run AI local

  • LM Studio. prettier GUI, same engine, friendly for non-technical users.
  • Jan.ai. open-source ChatGPT clone you self-host.
  • GPT4All. one-click on any laptop.
  • vLLM / TGI. production server stack for classified-area scope and regulated rooms.
  • Apple Intelligence. Apple’s local model on every M-series Mac and recent iPhone.
  • NVIDIA NIM / Chat with RTX. turnkey local AI on RTX Windows machines.
  • Home Assistant + Llama. voice control of the smart home, fully local, no Alexa or Google.

Datacenter vs local AI · two cost curves, one decision.

Datacenter / CloudLocal
Up-front cost$0$1,400 to $10,000 hardware
Recurring cost$20K to $60K/yr corporate, $20 to $60/mo personalPower only (~$15 to $80/mo)
PrivacyData leaves the buildingNothing leaves
Latency200 to 800ms50 to 150ms
Frontier gapBest models, day-zero6 to 12 months behind frontier
Internet outageDiesKeeps working
ComplianceVendor’s postureYours, end-to-end
ScaleInfinite, pay as you goCapped by your hardware
Update cycleVendor updates silentlyYou update on your schedule
Best forBleeding-edge tasks, big spikesDaily use, privacy-sensitive work

The honest split: 80% local, 20% cloud. Daily voice, transcription, camera analytics, summarization, smart-home brain. local. Bleeding-edge research and the occasional GPT-5 reasoning task. cloud. With proper segmentation, the user never knows which one answered.

The true cost to power AI · watts, water, silicon.

Datacenter side (the part nobody itemizes)

  • A single GPT-5 query uses ~3 watt-hours. About 10x a Google search.
  • A 1-minute Sora 2 clip burns ~3,000 to 5,000 Wh to generate. The energy of a microwave running for an hour.
  • A modern AI datacenter pulls 100 to 300 megawatts. The biggest under construction (Stargate, xAI Colossus, Microsoft Mt Pleasant) target 1 gigawatt. A small city.
  • Global datacenter power demand was ~1.5% of world electricity in 2022. Projected ~4% by 2027 (IEA). AI is 60 to 70% of that growth.
  • Water: a 100-token reply uses about half a liter of fresh water at most hyperscale sites.
  • Carbon: a single large training run emits 300 to 500 metric tons of CO2 (about 100 transatlantic flights).
  • The grid bill: northern Virginia residential rates went up ~12% in 2025 because of AI build-out. NJ and CT are next.

Local side (what your client actually pays)

  • Mac Studio M4 Max under realistic AI load: ~150 to 250W. About a desk lamp on steroids. ~$15 to $25/month at NJ rates.
  • NVIDIA RTX 5090 workstation: ~450 to 600W. ~$35 to $60/month.
  • DGX Spark or dedicated AI server: 800W to 2kW. ~$80 to $200/month.
  • Multiply by uptime, add ~30% cooling overhead, and you have the real number.

The hidden cost most integrators ignore

  • Heat in the rack. Local AI hardware drops 500 to 2,000W of heat into the AV closet 24/7. Plan a mini-split, not a fan.
  • UPS sizing. A real local AI rig needs a 1,500 to 3,000VA UPS. The standard rack UPS is undersized.
  • Power conditioning. GPUs hate dirty power. A managed PDU and surge protection are not optional.
  • Noise. Server-grade GPUs scream. Plan acoustic isolation if the rack lives near a bedroom or office.

A properly sized local AI box pays for itself faster than a luxury espresso machine and produces more business value.

Open source AI · the Linux moment for models.

Open weight vs true open source

  • Open weight (Llama 4, DeepSeek V3, Mistral, Qwen, Gemma): trained model is downloadable and runnable. Training data is usually NOT public. You can run, fine-tune, and deploy it.
  • True open source (OLMo, Pythia, Falcon, BLOOM): weights AND training data AND code are all public. Auditable end-to-end. Smaller and usually less capable, but you know everything about them.

The leaders right now

ModelMakerStrengthSize to run
Llama 4MetaBest general-purpose open70B
DeepSeek V3 / R1DeepSeek (China)Best reasoning, rivals GPT-570B / 685B MoE
Mistral Large 3Mistral (France)Fast, efficient, EU-aligned70B
Qwen 3AlibabaBest multilingual72B
Gemma 3GoogleSmall, fast, embedded27B
Phi-4MicrosoftTiny, runs on a laptop14B
Falcon 3TII (UAE)Apache 2.0, fully open40B
OLMo 2Allen InstituteTruly open, fully auditable32B

Why open source matters in our work

  • Privacy by default. Run locally, nothing leaves the building.
  • No vendor lock-in. Your client’s smart-home AI brain is portable.
  • No surprise pricing changes.
  • Air-gap compliant for government, military, hospital, financial services.
  • Customizable. Fine-tune Llama 4 on a family’s preferences, a hotel’s brand voice, or a boardroom’s standards.
  • Auditable. When a family asks "what did you put in our home", you can answer with specifics.

Watch the license

Llama 4 has a 700M-user clause that triggers a paid license at scale. DeepSeek and Qwen are made in China, relevant for some federal work. Mistral and Falcon are permissive Apache 2.0. Read the license. Have counsel sign off for regulated deployments.

Control systems that work with AI · the conductor matters.

Tier 1. production-ready AI integration

  • Crestron Home + Crestron Pro. Native Josh.ai driver, OpenAI/Anthropic via custom modules, AI camera analytics through Crestron 1Beyond. The only true commercial-grade brain that lives in a rack and survives. We are an Elite Pro Crestron Dealer. This is what we deploy.
  • Josh.ai. Purpose-built voice AI for luxury homes. Now ships with on-device processing. Integrates cleanly with Crestron, Lutron, Sonos, and most major brands. Our pick for high-end residential voice.
  • Control4 (Snap One). Halo Touch panels, Alexa and Google integration. Newer Halo models support local AI. Mid-tier residential, wider dealer network.
  • Savant. Savant Voice, Alexa and HomeKit integration, AI scene engine. Best for Apple-aligned homes.
  • Loxone. AI-driven heating curves, predictive shading, occupancy learning baked into the platform. Genuinely the best BMS and energy AI in the industry. Now expanding into US residential, commercial, and hospitality.
  • KNX. Open standard, AI lives in whatever gateway you put on top of it. Coming to America properly in 2026. Powerful for integrators willing to learn it.

Tier 2. workplace and conference AI

  • Microsoft Teams Rooms on Crestron, Logitech, Poly, Yealink. auto-framing, transcription, translation, Copilot integration.
  • Zoom Rooms. AI Companion, smart name tags, captions, meeting summary.
  • Cisco Webex / Cisco Rooms. strong AI for enterprise that lives in Cisco.
  • Q-SYS (QSC). AI auto-framing, beam steering, clean integration with major voice platforms.
  • Logitech CollabOS / Sight. AI camera that understands the room. Best price-performance.

Tier 3. voice and assistant layer

  • Apple HomeKit + Apple Intelligence. Local voice processing. Best privacy posture among consumer voice options.
  • Amazon Alexa+. Best ecosystem breadth. Worst privacy posture. Avoid for sensitive spaces.
  • Google Home + Gemini. Strong analytics. Mid-tier privacy. Tight with Nest cameras.
  • Home Assistant + Llama (open source). Fully local, fully open, fully customizable. Pair with Crestron or Loxone for serious systems.

Tier 4. camera and security AI

  • Verkada. cloud AI for video. Strong analytics, privacy is "trust us".
  • Avigilon (Motorola). best on-prem AI camera platform. Federal-friendly.
  • Hanwha Wisenet. good on-camera AI, fair price.
  • Ubiquiti UniFi Protect. local AI analytics, tight network integration. We deploy this on most homes.
  • Axis Communications. open API, partner-driven analytics, government-grade.

What we tell clients

  • Luxury home: Crestron + Josh.ai + Apple Intelligence + Llama 4 local. That is the stack.
  • Smart energy / BMS: Loxone or KNX layered under Crestron.
  • Boardroom / commercial: Crestron or Q-SYS as the brain, Teams or Zoom certified gear, plus a local AI box for transcription and summary.
  • Hospitality: Crestron for guest experience, Loxone for the BMS, Josh or Apple for voice, local AI server for analytics.
  • Government: Crestron, on-prem only, no cloud voice, no consumer assistants, Llama 4 or DeepSeek on approved hardware.

Security and privacy with AI · the real talk.

The seven privacy postures every AI device has

  1. What it captures. audio, video, occupancy, behavior, calendars, contacts.
  2. Where it processes. on-device, on a local hub, or in the vendor cloud.
  3. What it transmits. full audio, transcripts, just metadata, or nothing.
  4. Where it stores. for how long, in what jurisdiction, who can subpoena it.
  5. Who it shares with. third-party analytics, ad networks, training pipelines, law enforcement.
  6. Who can access it. vendor employees, contractors, partner cloud platforms.
  7. What happens when you cancel. does data delete, or sit in a backup forever.

If your integrator cannot answer all seven for every AI device they put in your home or building, do not let them put it in your home or building.

The eight real attack surfaces

  1. Always-on microphones. Echo, Google Home, Josh.ai, smart TV mics, conferencing rooms. Each one is a recording posture.
  2. AI cameras. Cloud-based VMS had real breaches in 2023 to 2025. Local-first is the answer.
  3. Voice cloning. Three seconds of your voice from a podcast or a Zoom is enough to clone you. Establish a family or company safe word.
  4. Prompt injection. Hidden text that tricks an AI agent into doing something the user never asked.
  5. Model exfiltration. Cloud AI can be jailbroken to leak training data.
  6. API key theft. Leaked keys get used to mine crypto or run scams on your bill. Cap every key.
  7. Default-credential AV gear. Devices left at admin/admin are scanned constantly.
  8. Network segmentation failure. AI cameras + IoT + control + work laptops on one flat VLAN is a free pivot.

The Restrepo privacy stack

  1. Local-first by default. Llama 4 / DeepSeek on a Mac Studio or NVIDIA box in the rack. Cloud is opt-in.
  2. Network segmentation. AI cameras, voice assistants, control processors, work devices, guest Wi-Fi. separate VLANs, separate policies, separate firewall posture. Ubiquiti UISP / Pro Partner work.
  3. Hardened credentials. Every device gets a unique strong credential, stored in our team password manager, rotated annually.
  4. Patched firmware. Every device on a documented update cadence.
  5. Privacy posture document. Every client gets a written summary of every AI device, what it captures, where the data goes, and how to revoke it.
  6. Voice-policy review. Where do you NOT want a microphone. We honor that. We document it.
  7. Recording retention. Cameras and voice get explicit retention windows, not "forever" by default.
  8. Family / staff safe-word policy. For high-net-worth and executive clients, a verbal safe word for any AI-related transaction request.
  9. Quarterly security review. Network, credentials, firmware, and privacy posture reviewed every 90 days.
  10. Documented exit plan. When you leave us or change platforms, we hand you the data, wipe what is wipeable, and document what is not.

The compliance lens

  • GDPR / CCPA / CPRA. Cloud AI usually fails this without a DPA.
  • HIPAA. No cloud AI without a BAA. Local-only is the safe answer.
  • CMMC, FedRAMP, FISMA. Cloud AI is mostly a non-starter. On-prem, approved-product list, sustainment plan.
  • NDAA Section 889 / TAA. Affects which AI cameras and AI gear are even buyable for federal work.
  • NJ Consumer Privacy Act (active 2025). Your home state. New disclosure requirements.
  • PCI-DSS. Hotels and restaurants. AI that touches the POS network is a PCI scope expansion.

Ten questions to ask any integrator before signing

  1. Where does the audio from this device go?
  2. What VLAN will the AI cameras live on?
  3. Are you giving me a written privacy posture document?
  4. Can the system run if the internet drops?
  5. Who has access to the cloud account behind this?
  6. What happens to my data if I cancel?
  7. Is the firmware patched and on a schedule?
  8. Are there any default passwords still in this system?
  9. Have you read the GDPR / CCPA / HIPAA implications for my use case?
  10. What is your incident-response plan if an AI device is compromised?

If they fumble three of those, walk away.

Building or refreshing a smart space?

If AI is on your roadmap for a luxury home, hospitality property, boardroom, or government facility. let’s walk it. We build local-first systems with a written privacy posture and a 10-year sustainment plan.