Why trust what the house publishes.
This is GenAI BR's charter: whom we write for, how an edition is made, what every number and every judgment means, and the scoreboard of what we got right and wrong. The page is versioned. When the method changes, it changes here before it changes in practice.
For whom
We write for the engineer who will put an agent into production on Tuesday, in Brazil.
She is a concrete person, not an audience. She has a system to ship, an inference budget, a committee asking "why not the open model?" and one hour a week to learn what actually changed. Cinco is written for that hour. If a paragraph does not change what she does on Tuesday, it goes.
Around her are two rings. Those who decide — purchases, hiring, architecture, regulation — read what she reads but need the result on one page: the second cut, Cinco for decision-makers, exists for that ring. And those starting out, who prefer five things explained to fifty headlines, and find the glossary and the trails in the Archive.
What the reader is not: an investor looking for a thesis, a journalist looking for a headline, a vendor looking for validation. They are welcome; the page is not written for them.
- The engineer
Builds with AI and needs to know what actually changed, not what was announced. The whole essay is for her: the thesis, the movements, the code.
- Those who decide
Buy, hire, design the architecture or regulate. They read the judgments and the panel; from October, they get the same week on one page.
- Those starting out
Prefer five explained items to fifty headlines. They come in through the glossary and the Archive's trails, and climb when they want.
How an edition is made
One reporting, two cuts.
The week starts with a pile of a couple of hundred things — papers, repositories, official documents, announcements — and ends with a thesis carried by five movements. The funnel below shows the stages and the day of each. The initial triage is AI-assisted; the reading, the checking, the code and the text are one person's.
Two hundred become five
The funnel of a typical week. The numbers are averages of the first weeks and will be updated with each edition.
- ~200skimmed
Titles, abstracts and metadata of everything that came out: arXiv, repositories, announcements, official gazette. AI-assisted triage flags what deserves reading.
- ~40actually read
Full paper, open code, whole document. Here the AI leaves the room.
- 12verified
Every number checked against the source. Self-reported marked. The arithmetic that carries the argument becomes code and is run.
- 5published
The movements that carry the week's thesis. Each primary source enters the Archive dated — and stays.
The filter
One question decides what gets in: does this change a builder's decision in the coming months? Five criteria hold it up.
- 1Primary source or nothing
Every item leads to the paper, the code, the official document or the original announcement. Press coverage enters as complementary reading, never as the source.
- 2Result before opinion
Each entry separates what was demonstrated from what was not. Self-reported numbers are marked as such.
- 3Relevance for builders
What changes an engineering, product or purchasing decision in the coming months gets in. What is only noise stays out, even if it is news.
- 4Brazil when there is a document
Regulation, public money and the labor market enter when there is an official text, a number or a deadline — not when there is only a promise.
- 5Conflicts declared
When an item touches something the house has a relationship with, it is written in the entry.
The two cuts
It is The Economist's rule: one reporting, several forms. What changes between the cuts is density, never the conclusion.
- CincoLive since September 2026
The essay. The thesis, five judgments with confidence levels, the panel, five movements with figures and executed code, what to watch, the numbered sources. Forty minutes of serious reading, every Tuesday at 7am.
- Cinco for decision-makersPlanned for October 2026
The same week on one page: what changed, what to do with it, what is still unproven. Nothing in it that was not reported in the essay; no judgment differs between the two.
Every edition has the same seven parts
On the left, the real edition no. 1. On the right, what each part promises. Repetition is what makes one week comparable to another — and accountable.
In 2023, the most capable model in the world charged US$ 60 per million output tokens. In April 2025 it charged US$ 8. Last Thursday, GPT-6 Astra shipped at US$ 50 — the same price as Claude Fable 5.1, and two and a half times the model OpenAI launched in June. In the same two-week window, Z.ai put a 320-billion-parameter model, MIT-licensed, served according to the company only on Chinese chips, at US$ 0.50. The distance between the two ends is a hundredfold. That is what we call the scissors: the blades move apart, and what sits between them is the entire market of people who build with AI.
1For 2026-class agentic work, cost per token no longer decides between closed and open. What decides is domain reliability, governance of the reasoning, and where the model runs.
100×The scissors — closed frontier ÷ permissive open
IIFifteen cents
Self-reported on the model card: 65.7% on SWE-bench Pro. There is no technical report. None of the four scores has been replicated.
Bids open for the Macaíba tender. Who shows up, with which GPU, and whether anyone contests the local-content clause.
1Hy4-preview — model card
Three paragraphs on what the week means, not what happened. Read only this and you leave knowing the house's position.
Five conclusions, each with a confidence level — high, moderate, low — in the vocabulary of intelligence analysis. The next edition revisits them.
Six fixed indicators, read again every week: frontier price, open price, the scissors, two benchmark ceilings, public money that became a tender.
Five chapters of the same argument. Each with a figure or an executed code demonstration — and the data table always one click away.
Self-reported or independent, on every number. Sample, interval, harness. What we do not know is written down.
Dates and conditions that would confirm or refute the judgments — so the next edition can be held to account.
Numbered, dated, with what we took from each. Press only when it is the sole source, and said so.
What the numbers mean
A number in this house is a number someone published, with the source next to it. Never an estimate of ours presented as data.
- Self-reported or independent
- Every result says who measured it. A number from the lab itself is marked self-reported; a replication is called one. The two never share a bar without warning.
- Bars always from zero
- We never truncate the axis of a bar chart. If the difference only shows with a cut axis, the difference is not the argument.
- One scale per chart, and the log declared
- Every figure has one scale. When it is logarithmic — prices, parameters — it says so in the subtitle, and the axis ticks show it.
- Direct labels, not a legend
- The name sits next to the point or the line. A scatter without a name on every point is a cloud, not an argument.
- Every figure has its table
- Every chart carries its data as a table, one click away, with the source. What cannot be read as a table does not become a chart.
- Executed code, not illustrative
- When arithmetic carries the argument, it is written as code, it was run, and the published output is the real output, dated. If you run it and get something else, write in.
- Sample and interval, when they exist
- A 119-task benchmark does not separate two models by two points; we say when a difference is noise. When the interval was not published, we say it was not.
- Numbers in Portuguese
- US$ 0,50, 47,9%, 7.000 petaflops. Decimal comma and thousands point, as in the country where the page is read.
The judgment vocabulary
Every edition closes with judgments: what we conclude and how sure we are, in the vocabulary of intelligence analysis (the ICD 203 directive, adapted).
- High confidence
- Independent sources converge, the mechanism is known and the track record is good. We would be surprised to be wrong.
- Moderate confidence
- The evidence is plausible and consistent but has a gap: a single source, a self-reported number, a short horizon. It is our reading, and it may change.
- Low confidence
- The evidence is fragmentary or the track record is poor. We publish because the question matters, not because the answer is firm.
What happens next
Every judgment is born open and gets, in "What to watch", the dates and conditions that would confirm or refute it. The edition that passes those dates comes back to it — confirmed, refuted or revised — with a line on what happened. We never delete a wrong judgment; the scoreboard below is the sum of all of them.
- open
- Published, waiting for the dates in "What to watch".
- confirmed
- The condition happened. It stays on the scoreboard as a hit, with the edition that closed it.
- refuted
- The opposite happened, or the deadline passed without the condition. It stays on the scoreboard as a miss, with the same visibility.
- revised
- We changed the wording or the confidence before the deadline. The earlier text stays visible.
Indicators
Six fixed series, read again every edition. Each reading is a number someone published, with the source. When an indicator's method changes, the series restarts and the earlier one stays. The whole series is here and as CSV.
Download the series as CSV↓Closed frontier price
How it is measured. List price, short context, of the most expensive general-purpose model among OpenAI, Anthropic and Google on the edition's day.
- no. 1$5008 Sept 2026 · GPT-6 Astra and Claude Fable 5.1 · openai.com
Edition Date Value Who Source no. 1 08 Sept 2026 $50 GPT-6 Astra and Claude Fable 5.1 openai.com Permissive open price
How it is measured. Lowest list price, on the lab's own API, among open-weight models under MIT or Apache 2.0 with an Artificial Analysis Intelligence Index ≥ 55.
- no. 1$0.5008 Sept 2026 · GLM-5.3-Flash (Z.ai) · huggingface.co
Edition Date Value Who Source no. 1 08 Sept 2026 $0.50 GLM-5.3-Flash (Z.ai) huggingface.co The scissors
How it is measured. Ratio between the two indicators above. It measures the distance in price, not in capability.
- no. 1100×08 Sept 2026 · 50 ÷ 0.50 · Closed frontier price
Edition Date Value Who Source no. 1 08 Sept 2026 100× 50 ÷ 0.50 Closed frontier price Terminal-Bench 4.0 ceiling
How it is measured. Highest published score on the current Terminal-Bench version, self-reported or independent — always stated which. When the version changes, the series restarts.
- no. 157.9%08 Sept 2026 · GPT-6 Astra, self-reported · openai.com
Edition Date Value Who Source no. 1 08 Sept 2026 57.9% GPT-6 Astra, self-reported openai.com SWE-bench Science ceiling
How it is measured. Best model on the leaderboard kept by the OpenMOSS group (Fudan), independent of vendors. 119 tasks: differences of a few points are noise.
- no. 147.9%08 Sept 2026 · Claude Opus 5, max effort · swescience.github.io
Edition Date Value Who Source no. 1 08 Sept 2026 47.9% Claude Opus 5, max effort swescience.github.io PBIA: money that became a tender
How it is measured. Sum of published tenders with a value and a deadline under Brazil's AI Plan (R$ 23 billion announced for 2024–2028). Announcements and partnerships without a tender do not count.
- no. 195908 Sept 2026 · Macaíba supercomputer (LNCC/FACC) · facc.org.br
Edition Date Value Who Source no. 1 08 Sept 2026 959 Macaíba supercomputer (LNCC/FACC) facc.org.br
Judgment scoreboard
- 5
- 5
- 0
- 0
- 0
Every judgment published, from every edition, with its state. Generated from the editions' content, not typed by hand.
For 2026-class agentic work, cost per token no longer decides between closed and open. What decides is domain reliability, governance of the reasoning, and where the model runs.
The ~100× gap between the closed frontier and permissive open models should persist for the next two quarters: closed labs are pricing reasoning and agency, not tokens, and open labs are pricing adoption.
SWE-bench Verified no longer discriminates between frontier models; small-sample domain leaderboards do not separate the top five either. Whoever decides a purchase needs their own evaluation.
Chain-of-thought monitorability will keep falling in flagships. “Reading the reasoning” should not be treated as a safety control in any production system.
The Macaíba supercomputer will probably be contracted in 2026; it is unlikely to operate by the announced date (end of 2027), given three postponements and the absence of operating costs in the tender.
| Edition | Judgment | Confidence | State |
|---|---|---|---|
| no. 1 | For 2026-class agentic work, cost per token no longer decides between closed and open. What decides is domain reliability, governance of the reasoning, and where the model runs. | moderate confidence | open |
| no. 1 | The ~100× gap between the closed frontier and permissive open models should persist for the next two quarters: closed labs are pricing reasoning and agency, not tokens, and open labs are pricing adoption. | moderate confidence | open |
| no. 1 | SWE-bench Verified no longer discriminates between frontier models; small-sample domain leaderboards do not separate the top five either. Whoever decides a purchase needs their own evaluation. | high confidence | open |
| no. 1 | Chain-of-thought monitorability will keep falling in flagships. “Reading the reasoning” should not be treated as a safety control in any production system. | high confidence | open |
| no. 1 | The Macaíba supercomputer will probably be contracted in 2026; it is unlikely to operate by the announced date (end of 2027), given three postponements and the absence of operating costs in the tender. | low confidence | open |
What we do not do
- We do not train models, sell courses or run events
- We read, check and write. The money comes from company briefings and, from 2027, from the Radar and State of AI Brasil. Reading Cinco is free.
- We do not publish without a primary source
- An item without a paper, code, document or original announcement does not get in, however newsworthy. Press is complementary reading, and said so.
- We do not forecast without a confidence level
- Every judgment says how sure we are and what would refute it. Opinion without that is not published as a judgment.
- We do not delete what we got wrong
- A refuted judgment stays on the scoreboard. A factual error is corrected in the entry with the earlier text visible, and listed in the corrections.
- We do not take disguised paid content
- Sponsorship of artifacts — Radar, State of AI — only without editorial influence and with the name on the cover. A paid item dressed as an entry, never.
- We do not let AI decide what matters
- It triages and organizes; the choice, the checking and the text are one person's. Below it says exactly where it comes in.
- We do not track who reads
- No third-party scripts. The only form asks for an e-mail and says where it goes.
- We do not publish more to look bigger
- One edition a week, the same format. When the house grows, it will be to read deeper.
AI policy
We use AI where it is good and say exactly where. If this changes, this page changes first.
- Triage
- Filter hundreds of items a week, extract metadata (authors, date, source type, code availability) and flag what deserves reading.
- Deduplication
- Group the same result published in three different places.
- Translation
- The English version of this site is AI-translated from Portuguese and reviewed by a person. Portuguese is the original.
- Code
- This site was designed and built with an AI assistant, under human direction and review, decision by decision. It is described in "How this site was made", on the About page.
- Choose the movements
- Selection is human. AI suggests reading; it does not decide what matters.
- Write the final text
- No published paragraph is machine-generated. What you read in an edition or an entry was written by the person who read the source.
- Verify numbers
- Every number is checked by a person against the primary source. AI gets numbers wrong with confidence; that is why it does not do this part.
- Sign
- Every edition carries a signature. It is a person, with a name, answering for what is written.
Corrections
Errors are corrected in the entry or the edition, with the date and the earlier text visible, and listed here. An error that changes a movement's conclusion gets a note in the following edition.
No corrections recorded so far. The first will appear here, dated — and it will probably exist.
Versions
Every change to the method gets a line here, dated. A change in practice without a change here is an error of ours — and belongs in the corrections.
- v1
First version. Gathers what was spread across the About page — how we choose, the anatomy of an edition, the AI policy, the corrections — and adds the reader, the two cuts, the judgment vocabulary, the rules for numbers, the scoreboard and the indicator series.
Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.