Paper deep dive
Reading Copom's Tone: A Weighted LLM Framework for Hawkish-Dovish Sentiment, Forward Guidance, and Uncertainty
Gabriel de Macedo Santos
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper documents an applied natural-language-processing framework for measuring the tone of Brazilian Monetary Policy Committee (Copom) statements. The project is explicitly inspired by iSent, Itaú's Central Bank sentiment classifier, particularly its sentence-level division of official communication into hawkish, dovish, neutral, and out-of-context classes. The implementation extends that idea in three directions. First, an LLM identifies short hawkish and dovish expressions and assigns each a 0-to-1 intensity weight. Second, the document index combines sentence counts with document-specific average signal intensities, producing a bounded score from -1 to 1. Third, a separate full-document layer measures forward-guidance direction, guidance explicitness, uncertainty level, and change in uncertainty. The empirical sample is restricted to communications dated August 2016 or later and contains 80 statements and 1,498 classified sentences from August 31, 2016 through August 5, 2026. Across this sample, 33.3% of sentences are hawkish, 18.0% dovish, 42.1% neutral, and 6.5% out of context. The average document score is +0.107, while the most hawkish reading is +0.570 in August 2021. The latest statement, dated August 5, 2026, scores +0.232, with eight hawkish, two dovish, and nine neutral sentences. Its structural overlay is more nuanced: guidance is directionally ambiguous but partly explicit, while uncertainty is classified as central and higher than at the prior meeting. Tone and the guidance-direction score have a contemporaneous Pearson correlation of 0.719. These are descriptive outputs, not a validated forecast of Selic decisions or DI returns. The main contribution is therefore methodological: a transparent, incremental, auditable system that separates rhetorical tone from policy guidance and uncertainty.
Tags
Links
- Source: https://arxiv.org/abs/2608.07251v1
- Canonical: https://arxiv.org/abs/2608.07251v1
Trouble viewing inline? Open PDF directly →
Full Text
35,210 characters extracted from source content.
Expand or collapse full text
1 Reading Copom's Tone: A Weighted LLM Framework for Hawkish-Dovish Sentiment, Forward Guidance, and Uncertainty Sample used in this version: 80 policy statements, August 2016 to August 2026 Gabriel de Macedo Santos Instituto de Tecnologia e Liderança Abstract. This pa per documents a n applied na tura l-la ngua ge-processing framework for mea suring the tone of Bra zilia n Moneta ry Policy Committee (Copom) sta tements. The project is explicitly inspired by iSent, Itaú's Centra l Bank sentiment cla ssifier, pa rticula rly its sentence-level division of officia l communica tion into ha wkish, dovish, neutra l, a nd out-of-context cla sses. The implementation extends tha t idea in three directions. First, a n LLM identifies short ha wkish and dovish expressions and a ssigns ea ch a 0-to-1 intensity weight. Second, the document index combines sentence counts with document-specific a vera ge signa l intensities, producing a bounded score from -1 to 1. Third, a sepa ra te full-document la yer mea sures forwa rd-guida nce direction, guida nce explicitness, uncerta inty level, and change in uncerta inty. The empirica l sa mple is restricted to communica tions dated August 2016 or la ter a nd conta ins 80 sta tements and 1,498 cla ssified sentences from August 31, 2016 through August 5, 2026. Across this sample, 33.3% of sentences a re ha wkish, 18.0% dovish, 42.1% neutra l, a nd 6.5% out of context. The avera ge document score is +0.107, while the most ha wkish reading is +0.570 in August 2021. The la test sta tement, dated August 5, 2026, scores +0.232, with eight ha wkish, two dovish, a nd nine neutra l sentences. Its structura l overlay is more nuanced: guidance is directiona lly a mbiguous but pa rtly explicit, while uncerta inty is cla ssified a s centra l a nd higher than a t the prior meeting. Tone and the guida nce-direction score have a contempora neous Pea rson correla tion of 0.719. These a re descriptive outputs, not a va lida ted foreca st of Selic decisions or DI returns. The ma in contribution is therefore methodologica l: a transpa rent, incrementa l, audita ble system tha t sepa ra tes rhetorica l tone from policy guida nce a nd uncerta inty. Keywords: central bank communication; Copom; sentiment analysis; large language models; forward guidance; monetary policy; Brazil. JEL Classification: E52, E58, C45, C55, G14. 1. Introduction Centra l ba nk communication is not a ncilla ry to moneta ry policy. It shapes expecta tions about the rea ction function, the likely pa th of short-term interest rates, the ba la nce of risks, a nd the conditions under which policyma kers may cha nge course. In infla tion-ta rgeting regimes, the la ngua ge surrounding a ra te decision ca n therefore ca rry information that is distinct from the mecha nica l decision itself. Blinder et a l. (2008) describe communica tion as a potentia lly powerful pa rt of the centra l bank's toolkit, while Ha nsen a nd McMahon (2016) show that communica tion about forwa rd guidance can ma tter more for ma rkets tha n descriptions of current economic conditions. The pra ctica l difficulty is mea surement. Copom sta tements conta in simulta neous signa ls a bout infla tion, a ctivity, the exchange rate, fisca l policy, externa l risks, expecta tions, and the degree of confidence in the ba seline scena rio. A single statement may describe ea sing in current a ctivity while wa rning a bout unanchored infla tion expecta tions. A pure word count ca n miss this context, while a document-level LLM la bel ca n hide interna l disa greement a cross sentences. This project a ddresses tha t problem with a sentence-level, prompt-ba sed cla ssifier a nd a sepa ra te structura l overlay. Its sta rting inspira tion is iSent, introduced by Itaú Uniba nco Ma cro Resea rch in July 2024. iSent cla ssifies sentences from officia l centra l-ba nk documents a s dovish, neutra l, ha wkish, or out of context, a nd a ggrega tes the rela tive presence of those cla sses into a n index between -1 and 1. The current project a dopts that intuitive unit of a na lysis a nd cla s ta xonomy, while delibera tely cha nging the ca libra tion a nd output a rchitecture. 2 The framework adds phra se-level intensity weights, trea ts forma l decision-only sentences a s neutra l by rule, and estima tes forwa rd guida nce a nd uncerta inty independently of the tone score. It a lso uses a n incrementa l ca che so tha t new statements ca n be processed without re-annotating the historica l corpus. These a dditions a re useful beca use ha wkish tone, explicit guida nce towa rd a hike, a nd high uncerta inty a re related but not identica l objects. The latest rea ding illustra tes the distinction: the August 2026 sta tement is ha wkish in tone, but its policy guida nce is directiona lly a mbiguous. The pa per makes three contributions. First, it documents the exa ct scoring formula a nd the data linea ge from sentence annota tions to document scores. Second, it presents a consistently defined historica l sample from August 2016 to August 2026, excluding a l ea rlier observations from the empirica l a na lysis. Third, it provides a n implementation a udit that sepa rates wha t is currently operationa l from wha t rema ins a va lidation or resea rch extension. The result is best viewed a s a monitoring index, not yet a s a tra ding signa l or ca usa l estima te of communica tion effects. 2. Related Literature and the iSent Inspiration The empirica l litera ture on centra l-bank text ha s moved from dictiona ries a nd ba g-of-words representa tions towa rd contextua l la ngua ge models. Correa et a l. (2021) construct a doma in-specific dictiona ry for fina ncia l-stability reports, illustra ting the va lue of ta iloring sentiment vocabula ries to policy la ngua ge. Hansen a nd McMahon (2016) use computa tiona l linguistics to sepa ra te communica tion about economic conditions from forwa rd guidance. More recent work uses models tra ined or a dapted to the centra l-ba nking doma in, including Centra lBa nkRoBERTa (Pfeifer and Ma rohl, 2023) a nd the CB-LMs introduced by Ga mba corta et a l. (2024). LLMs reduce some of the rigidity of dictiona ries by interpreting context, negation, conditiona ls, a nd policy-specific phra sing. They a lso create new risks: model outputs ca n va ry with prompts and versions; pretra ining data may conta in the documents being cla ssified; a nd reproducibility depends on pinning the model a nd the inference configura tion. Silva , Moriya , and Veyrune (2025) show the potentia l of sentence-level LLM cla ssification a t la rge sca le, but their framework a lso underscores the importa nce of sepa ra ting communication dimensions ra ther than colla psing a l la ngua ge into one sentiment la bel. The closest direct inspiration is Itaú's iSent. Its published methodology uses roughly one thousand economist-la beled Bra zilia n sentences, sentence-level GPT-4 c la ssifica tion, and retrieved simila r examples through a FAISS index. The iSent index is the difference between ha wkish and dovish sentence counts divided by the number of ha wkish, dovish, and neutra l sentences. Itaú reports correla tions of 0.79 with the contempora neous Selic cha nge a nd 0.77 with the one- meeting-a head cha nge for Bra zil. Those figures a re properties of iSent a nd should not be attributed to the index developed here. Dimension iSent (Itaú, 2024) This project Unit of analysis Sentence Sentence for tone; full document for structural dimensions Tone classes Hawkish, dovish, neutral, out of context Hawk, dove, neutral, out Model grounding Economist-labeled examples; GPT-4; FAISS retrieval Prompt instructions and three heuristic anchors; API-hosted LLM Index Relative class presence; unweighted Class counts scaled by document-specific mean phrase intensity Additional dimensions Published policy/market comparison and rates backtest Guidance direction, explicitness, uncertainty level, uncertainty change Retrieval in baseline Yes No; the current FAISS routine is optional and not called by sentence classification Current validation status Published correlation and backtest Descriptive implementation; external validation still required Table 1. Methodological relationship between iSent and the current project. Note: The project is independently developed and is not affiliated with or endorsed by Itaú Unibanco. iSent is cited as methodological inspiration. 3 3. Data and Project Architecture The empirica l sa mple covers 80 Copom statement da tes from August 31, 2016 through August 5, 2026. The August 2016 cutoff is applied before computing any descriptive statistic, sentence sha re, correlation, extreme rea ding, or regime avera ge in this pa per. The prima ry source is the officia l Portuguese text of Copom sta tements. The code expects one record per da te, removes rows with missing da tes or text, sorts chronologica lly, a nd reta ins the la st observation when duplicate dates exist. The supplied ana lytica l files conta in 1,498 cla ssified sentence objects within the reta ined window. Of those, 1,400 enter the score denomina tor; 98 out-of-context sentences a re excluded. The ra w sta tements file a nd externa l model configura tion a re not included in the supplied a rtifact set. The a nnota tions, scores, structura l outputs, and code a re sufficient to reproduce the descriptive results in this paper, but strict end-to-end replica tion would a lso require the ra w sta tements, model identifier, prompt version, a nd inference settings. This distinction is recorded in the implementa tion a udit ra ther tha n obscured. Artifact Observed content Role in the framework sentiment_classifier.py 613-line incremental pipeline Sentence splitting, LLM prompts, scoring, structural classification, plot, optional FAISS sentence_annotations.json 80 retained dates; sentence labels and phrase signals Primary audit trail after the August 2016 sample filter scores_comunicados.json 80 retained document scores and class counts Weighted tone index for the study window structural_analysis.json 80 retained document-level analyses Guidance direction/explicitness and uncertainty level/change pesos_hawk_dove.json Full-file diagnostic weights Not used in sample statistics; scores use document-specific means text_features_clean.csv 78 statements, Aug. 2016-Apr. 2026 Convenience table; JSON adds June and August 2026 hawk_dove_history.png History plot from Aug. 2016 Published visualization supplied with the project Table 2. Uploaded artifacts and their economic or implementation role. Source: Uploaded project files. Counts were recomputed from the supplied JSON and CSV outputs. Figure 1. Project pipeline from official text to the tone index and structural overlay. Source: Author's implementation based on the uploaded classifier. 4. Methodology 4.1. Text normalization and sentence segmentation The pipeline first converts the sta tements into a two-column date-text da ta set. Dates a re pa rsed, inva lid observations a re dropped, and duplica te da tes a re deduplica ted. New lines a re repla ced with spa ces before cla ssifica tion. Sentence bounda ries a re identified with a regula r expression that looks for whitespace following periods, question ma rks, exclamation points, semicolons, or colons when the next token begins with a n upperca se letter, number, quotation ma rk, or pa renthesis. Segments with fewer tha n four whitespa ce-delimited tokens a re excluded. 4 This deterministic splitter is efficient, but it inherits formatting defects from scraped source text. Severa l a rchived annota tions conta in missing spa ces after punctuation, causing logica lly sepa ra te sentences to be merged. This matters beca use a merged segment ca n conta in both ha wkish a nd dovish phra ses while receiv ing only one fina l cla s. The structura l full-document la yer is less exposed to this segmenta tion issue, but the tone score is not. 4.2. Sentence classification and phrase-level signals Ea ch segment is sent to an LLM under a Portuguese prompt that defines four mutua lly exclusive cla sses. Ha wkish sentences empha size persistent or eleva ted inflation, upside risks, restrictive policy, higher ra tes, fisca l deteriora tion, or other forces that increa se the required degree of moneta ry restra int. Dovish sentences empha size disinfla tion, economic sla ck, wea k a ctivity, ea sing, or prospective ra te cuts. Neutra l sentences a re informa tiona l or ba la nced. Out-of-context sentences include a dministra tive content, na mes, a nd ma teria l outside moneta ry policy or ma croeconomic conditions. The prompt a lso a sks the model to extract short origina l expressions that support a ha wkish or dovish rea ding. Each signa l is a ssigned a weight from 0 to 1: 0.1-0.3 for wea k, 0.4-0.6 for modera te, a nd 0.7-1.0 for strong evidence. Weights and labels a re va lida ted a nd clipped after pa rsing. The model can return signa ls of both signs inside one segment, even though the segment ha s one fina l cla s. Up to forty segments a re cla ssified per API ca l. A specia l instruction trea ts a sentence that only reports the mecha nica l ra te decision a s neutra l with no signa ls. This is economica lly motivated: the index is intended to ca pture the communication surrounding the a ction, not mecha nica lly encode the ra te cha nge twice. The prompt further provides specia l guidance for the ba la nce of infla tion risks and includes three historica l ca libra tion a nchors: Ma rch 2020 a s strongly dovish, Ma y 2021 a s ha wkish, a nd August 2023 a s dovish. These a nchors guide the LLM; they a re not ha rd constra ints on the fina l score, which is recomputed mechanica lly from the returned la bels a nd weights. 4.3. Weighted document score For sta tement t, let 푁 퐻,푡 , 푁 퐷_푡 , a nd 푁 푁 denote the numbers of ha wkish, dovish, and neutra l sentences. Out-of-context sentences do not enter the denomina tor. Let the a vera ge intensity of a l ha wkish signa ls extra cted a nywhere in the document be 푤|퐻 푡 , a nd define 푤|퐷 푡 a na logously. The implementation uses a defa ult intensity of 1 when no signa l of a given sign is a va ila ble. The ra w score is: 푆푐표푟푒 푡 = 푤 | 퐻 푡 ×푁 퐻,푡 −푤 | 퐷 푡 ×푁 퐷,푡 푁 퐻 +푁 퐷 +푁 푁 The score is then clipped to the interva l [-1, 1]. Posit ive va lues indicate a net ha wkish tone a nd nega tive va lues a net dovish tone. Neutra l sentences dilute the ma gnitude without cha nging the numera tor. A subtle but important implementa tion deta il is that intensity avera ges a re computed from a l extra cted signa ls, not only from signa ls found in sentences whose fina l cla s ma tches the signa l. Thus, a neutra l or ha wkish segment that conta ins a dovish phra se ca n a ffect the dovish intensity term. Within the reta ined August 2016-August 2026 sample, ha wkish and dovish signa l weights a vera ge 0.654 and 0.586. These va lues a re dia gnostics and a re not used by compute_scores. The live document score instea d reca lcula tes both intensities within ea ch statement. This makes the index sensitive to loca l rhetorica l strength but a lso increa ses sampling noise in short documents. 4.4. Structural communication layer A second LLM ca l eva lua tes the entire statement a long four dimensions. Guidance direction equa ls -1 when the text points to ea sing, 0 when the next move is unclea r, a nd +1 when it points to tightening. Guida nce explicitness equa ls 0 for no guida nce, 0.5 for indirect or conditiona l guida nce, and 1 for explicit guida nce. Uncerta inty level ra nges from 0 to 3, while uncerta inty cha nge ta kes va lues -1, 0, or +1 rela tive to the prior meeting. GuidanceScore t = Direction t x Explicitness t The product preserves direction while shrinking indirect signa ls towa rd zero. It is not pa rt of the tone index. This sepa ration a llows the fra mework to identify statements tha t a re ha wkish in dia gnosis but neutra l a bout the next policy a ction, or dovish in guida nce while empha sizing unusua lly high uncerta inty. 5 4.5. Incremental execution and optional retrieval The pipeline is incrementa l. Existing a nnotations a nd structura l cla ssifica tions a re loa ded from JSON ca ches, a nd only da tes not a lrea dy present a re sent to the LLM. The complete a rchive is then rescored so tha t a l outputs rema in synchronized. A force fla g disca rds the ca che a nd re-annota tes the full sample. This a rchitecture reduces cost and la tency, but ca ched results a lso mea n that historica l cla ssifica tions ca n reflect different model versions unless model and prompt meta da ta a re stored with every run. The code can optiona lly build a FAISS vector index over full sta tement texts. In the supplied implementa tion, however, the sentence-cla ssifica tion function does not query tha t index or inject retrieved examples into the prompt. The baseline should therefore be described a s prompt-ba sed LLM cla ssification with optiona l retrieva l infra structure, not a s a n opera tiona l RAG cla ssifier. 5. Results 5.1. Corpus composition and signal intensities The August 2016-August 2026 sample conta ins 1,498 cla ssified segments. Neutra l sentences a re the la rgest cla s, a t 631 observations or 42.1% of the corpus. Ha wkish sentences tota l 499 (33.3%), dovish sentences 270 (18.0%), a nd out-of- context sentences 98 (6.5%). The preva lence of neutra l la ngua ge is economica lly pla usible beca use Copom sta tements conta in projections, descriptions, decision mecha nics, and a dministrative ma teria l that should not a l be ma pped into a policy sta nce. Figure 2. Sentence-level class composition, August 2016-August 2026. Note: Out-of-context sentences are excluded from the document-score denominator. The LLM extra cted 1,552 ha wkish phra se signa ls a nd 787 dovish signa ls within the reta ined sa mple. Their mea n weights a re 0.654 and 0.586, respectively, with sta nda rd devia tions close to 0.20. Signa l counts exceed sentence counts because one segment ca n conta in multiple releva nt phra ses. The difference in avera ge intensity contributes to a structura l positive tilt in the numera tor, but the a ctua l score uses document-specific ra ther tha n globa l a vera ges. Metric Aug. 2016-Aug. 2026 value Interpretation Statements 80 August 31, 2016-August 5, 2026 Classified segments 1,498 Includes out-of-context segments Segments in denominator 1,400 Hawk + dove + neutral 6 Metric Aug. 2016-Aug. 2026 value Interpretation Mean document score +0.107 Positive/hawkish tilt on average Median document score +0.106 Typical statement is mildly hawkish Standard deviation 0.208 Substantial meeting-to-meeting variation Interquartile range -0.041 to +0.265 Middle 50% of document scores Negative / nonnegative scores 26 / 54 Sign classification used by the history plot Table 3. Descriptive statistics of the weighted Copom tone index. 5.2. Historical path and monetary-policy regimes The reta ined history shows a clea r cyclica l pattern. From August 2016 through 2020, communica tion is dovish on avera ge (-0.078), consistent with repeated disinfla tion langua ge, sla ck, a nd ea sing guida nce. The series turns sha rply positive in 2021-2023, with a n a vera ge of +0.263, a s infla tion persistence, risk a symmetry, and the need for a bove- neutra l policy domina te the communica tion. The 2024-August 2026 window rema ins strongly positive at +0.237, even though its a vera ge guidance score is nea r zero. Tha t combina tion reflects a communica tion regime in which the dia gnosis ca n rema in restrictive while the immedia te direction of policy becomes conditiona l. Figure 3. Weighted hawkish-dovish score for the 80-statement study sample. Note: Positive readings are hawkish; negative readings are dovish. All observations before August 2016. 7 Figure 4. Mean document score by broad communication regime. Note: Regime boundaries are descriptive groupings selected for this paper, not estimated breakpoints. Period Documents Mean tone Mean guidance Mean uncertainty Aug. 2016-2020 35 -0.078 -0.457 2.23 2021-2023 24 +0.263 +0.458 2.54 2024-Aug. 2026 21 +0.237 -0.024 2.76 Table 4. Era-level averages of tone, guidance, and uncertainty. Note: Guidance equals direction multiplied by explicitness. Uncertainty ranges from 0 to 3. 5.3. Extreme readings The most dovish observa tion in the reta ined sample is Ja nua ry 11, 2017, a t -0.357, with two ha wkish, nine dovish, and two neutra l sentences. The rea ding reflects a n explicit intensifica tion of the ea sing cycle amid wea k a ctivity a nd ongoing disinfla tion. The strongest ha wkish reading is August 4, 2021, a t +0.570, supported by fourteen ha wkish sentences and no dovish sentences. Tha t statement empha sizes persistent infla tion, a n unfavorable composition of prices, eleva ted fisca l risk, a n upside-skewed ba la nce of risks, a nd a n explicit expecta tion of a nother a djustment of the sa me ma gnitude. The la rgest positive readings cluster a round the 2021 tightening cycle and the renewed ha wkish communica tion of late 2024 a nd ea rly 2025. The la rgest nega tive rea dings cluster a round the 2016-2018 ea sing cycle, when disinfla tion, economic slack, a nd forwa rd guida nce towa rd continued cuts were recurrent. These patterns a re economica lly coherent, but coherence is not the sa me a s out-of-sa mple predictive va lida tion. Date Score Hawk Dove Neutral Reading 2017-01-11 -0.357 2 9 2 Most dovish; easing intensification 2017-07-26 -0.300 0 7 6 Broadly dovish easing language 2021-08-04 +0.570 14 0 4 Most hawkish 2021-06-16 +0.487 13 1 3 Tightening cycle 2025-01-29 +0.479 11 0 6 Persistent inflation risks 2026-08-05 +0.232 8 2 9 Latest reading Table 5. Selected extreme and recent document scores. 8 Source: Author's calculations from the uploaded score and annotation files. 5.4. Tone, forward guidance, and uncertainty Guida nce direction is dovish for 32 statements, neutra l or ambiguous for 22, a nd ha wkish for 26. Guidance is fully explicit in 44 sta tements and pa rtia l or conditiona l in 36; none of the reta ined sta tements is cla ssified a s ha ving no guida nce. Uncerta inty is re levant or centra l (levels 2 or 3) in a l 80 statements. The uncerta inty-change va riable is mostly zero: only 13 increa ses and three decrea ses a re recorded, which pa rtly reflects the prompt instruction to use zero when compa rison with the prior meeting is not supported. The contempora neous Pea rson correla tion between the tone score a nd guida nce direction multiplied by explicitness is 0.719; the Spea rman correla tion is 0.718. This is a strong but incomplete a ssocia tion. It confirms tha t tone and guidance often move together, while lea ving substantia l room for divergence. Because both outputs a re genera ted by LLM prompts a pplied to the same documents, this correla tion is descriptive a nd should not be interpreted a s independent va lida tion. Figure 5. Tone, forward guidance, and uncertainty across Copom statements. Note: Tone and guidance use three-document moving averages for readability. Uncertainty is shown at the document level. 5.5. Latest reading: August 5, 2026 The latest statement scores +0.232. The cla ssifier identifies eight ha wkish, two dovish, and nine neutra l sentences. Ha wkish content includes an upside-skewed ba lance of infla tion risks, unanchored expectations, a resilient la bor ma rket, services infla tion risk, excha nge-rate depreciation risk, and the possibility that demand stimulus wea kens moneta ry transmission. Dovish content includes the gra dua l modera tion of a ctivity, disinfla tion in hea dline a nd underlying mea sures, a nd downside scena rios tied to a sha rper domestic or globa l slowdown. The structura l la yer a ssigns guidance direction 0 and explicitness 0.5. In other words, the statement conta ins conditiona l guida nce but does not clea rly commit to either continued cuts or renewed tightening. Uncerta inty is level 3 a nd cha nge +1, reflecting the centra lity of externa l conflict, a sset a nd commodity vola tility, fisca l concerns, a nd the statement's explicit description of a significa nt increa se in uncerta inty. The economic interpreta tion is therefore not simply 'ha wkish.' It is a ha wkish risk dia gnosis combined with conditiona l, directiona lly a mbiguous next-step guida nce. Latest metric August 5, 2026 value Interpretation 9 Latest metric August 5, 2026 value Interpretation Weighted tone score +0.232 Net hawkish diagnosis Sentence mix 8 hawk / 2 dove / 9 neutral Hawkish content dominates directional sentences Guidance direction 0 No clear next-move direction Guidance explicitness 0.5 Conditional or partial guidance Uncertainty level 3 Uncertainty is a central axis of the communication Uncertainty change +1 Higher than at the prior meeting Table 6. Latest Copom communication reading in the study sample. Source: Uploaded scores, sentence annotations, and structural analysis. 6. Validation, Robustness, and Limitations The strongest feature of the project is a uditability. Every score ca n be traced to sentence la bels, extra cted expressions, and weights. The code norma lizes la bels, clips weights, excludes out-of-context segments, and records structura l justifica tions. The incrementa l ca che ma kes the system opera tiona l for recurring Copom monitoring. The principa l limita tion is the a bsence of a manua lly la beled holdout set. The same LLM tha t interprets the text supplies both the cla s and the signa l weight. Without economist annotations, confusion ma trices, ca libration curves, or a greement statistics, the model's a ccura cy cannot be quantified. The three prompt anchors provide qua lita tive orienta tion, but their ta rget va lues do not constra in the fina l mechanica l score a nd the rea lized scores differ materia lly from the a nchor suggestions. A second limita tion is segmenta tion. Missing spa ces a fter punctuation can merge independent cla uses or sentences. Beca use one merged segment receives one cla s, the output ca n understate mixed communication. A dedica ted Portuguese sentence tokenizer, source-text clea ning, a nd unit tests a round a bbreviations, ta bles, a nd enumera ted risk lists would reduce this problem. A third limita tion concerns intensity a ggrega tion. The score multiplies a sentence count by the mea n weight of every same-sign signa l in the document. This is not equiva lent to summing sentence-level we ighted proba bilities. It gives identica l count weight to a strongly ha wkish sentence a nd a wea kly ha wkish sentence while sca ling the entire ha wkish count by the document's mea n signa l intensity. Short documents a nd documents with ma ny signa ls in one sentence can therefore beha ve differently from a n intuitive a dditive scheme. An unweighted iSent-style index a nd a sentence-level weighted sum should be reported a s robustness a lterna tives. Fina lly, the project does not yet test whether the score predicts Selic cha nges, DI returns, yield-curve shifts, or foreca st errors. Ita ú's reported correlations a nd ba cktest belong to iSent. Replica ting tha t va lida tion here would require merging the score with the exa ct decision cha nge, the one-meeting-ahea d change, and properly la gged ma rket prices, while preventing look-a hea d bia s. Tra ding results should include execution timing, contract construction, vola tility sca ling, tra nsa ction costs, a nd dra wdown dia gnostics. Robustness test Purpose Recommended output Economist-labeled gold set Measure classification validity Precision, recall, F1, confusion matrix, Krippendorff alpha Repeated LLM runs Measure stochastic instability Label agreement and score dispersion by statement Alternative model/prompt Assess model dependence Rank correlation and extreme-reading stability Unweighted index Benchmark against iSent-style aggregation Difference from weighted baseline Sentence-level weighted sum Test current mean-weight aggregation Level and turning-point comparison 10 Robustness test Purpose Recommended output Tokenizer ablation Quantify segmentation sensitivity Reclassified segments and score revisions Selic lead/lag study Test policy alignment Contemporaneous and one-meeting-ahead correlations DI event study/backtest Test market relevance Returns, information ratio, costs, drawdowns, OOS split Table 7. Priority validation and robustness program. Note: These tests are recommendations; their results are not present in the uploaded artifacts. 7. Conclusion This paper documents a weighted LLM framework for tra cking the ha wkish-dovish tone of Copom sta tements. Inspired by Itaú's iSent cla ssifier, the project uses sentence-level ha wk, dove, neutra l, a nd out la bels, but extends the a rchitecture with phra se intensity weights a nd a sepa ra te ana lysis of forwa rd guida nce and uncerta inty. The result is a n incrementa l monitoring pipeline whose outputs ca n be tra ced ba ck to individua l sentences a nd supporting expressions. Across 80 statements from August 2016 to August 2026, the avera ge tone score is +0.107. Communica tion is dovish on avera ge from August 2016 through 2020, sha rply ha wkish in 2021-2023, a nd rema ins positive in 2024-2026. The latest score is +0.232, but the structura l layer shows why a single la bel is insufficient: forwa rd guida nce is conditiona l and directiona lly a mbiguous, while uncerta inty is centra l a nd rising. The project's most credible current contribution is mea surement a nd orga niza tion, not prediction. It provides a disciplined wa y to sepa ra te the tone of the economic dia gnosis from the direction a nd explicitness of future policy guida nce. Converting this monitoring system into a publisha ble predictive index requires a huma n-la beled benchma rk, model-version controls, segmentation improvements, a ggrega tion a bla tions, and a genuinely out-of-sample Selic a nd DI va lida tion exercise. References Banco Centra l do Bra sil. (n.d.). Copom statements: Chronological archive. Retrieved August 6, 2026, from https://w.bcb.gov.br/en/moneta rypolicy/copomsta tements/cronologicos Blinder, A. S., Ehrma n, M., Fra tzscher, M., De Haa n, J., & Ja nsen, D.-J. (2008). Centra l bank communica tion and moneta ry policy: A survey of theory and evidence. Journal of Economic Literature, 46(4), 910–945. https://doi.org/10.1257/jel.46.4.910 Brown, T. B., Ma n, B., Ryder, N., Subbia h, M., Ka plan, J. D., Dha riwa l, P., Nee la kantan, A., Shyam, P., Sa stry, G., Askell, A., Aga rwa l, S., Herbert-Voss, A., Krueger, G., Henigha n, T., Child, R., Ramesh, A., Ziegle r, D. M., Wu, J., Winter, C., . . . Amodei, D. (2020). La ngua ge models a re few-shot lea rners. Advances in Neural Information Processing Systems, 33, 1877–1901. https://papers.nips.c/paper/2020/ha sh/1457c0d6bfcb4967418bfb8a c142f64a-Abstra ct.html Correa, R., Garud, K., Londono, J. M., & Mislang, N. (2021). Sentiment in central banks’ financial stability reports. Review of Finance, 25(1), 85–120. https://doi.org/10.1093/rof/rfa a 014 Ferreira , L. N., Ga rzeri, C., Guillen, D., Lima , A., & Monteiro, V. (2025). The not so quiet revolution: Signal and noise in central bank communication (Working Pa per Series No. 635). Ba nco Centra l do Bra sil. https://w.bcb.gov.br/content/publica coes/WorkingPa perSeries/WP635.pdf Ga mba corta , L., Kwon, B., Pa rk, T., Pa telli, P., & Zhu, S. (2024). CB-LMs: Language models for central banking (BIS Working Pa pers No. 1215). Ba nk for Interna tiona l Settlements. https://w.bis.org/publ/work1215.htm Ha nsen, S., & McMa hon, M. (2016). Shocking la ngua ge: Understanding the ma croeconomic effects of centra l bank communication. Journal of International Economics, 99(S1), S114–S133. https://doi.org/10.1016/j.jinteco.2015.12.008 Itaú Unibanco Ma cro Resea rch. (2024, July 5). iSent: Itaú’s central bank sentiment classifier [Ma cro Vision report]. https://macroa tta chment.cloud.itau.com.br/atta chments/fe3890ca-628a-4b80-a 41d- 49358d24046d/20240705_MACRO_VISION_iSent.pdf 11 Johnson, J., Douze, M., & Jé gou, H. (2017). Billion-scale similarity search with GPUs [Preprint]. a rXiv. https://doi.org/10.48550/a rXiv.1702.08734 Lewis, P., Perez, E., Piktus, A., Petroni, F., Ka rpukhin, V., Goya l, N., Küttler, H., Lewis, M., Yih, W.-T., Rocktä schel, T., Riedel, S., & Kiela , D. (2020). Retrieva l-augmented genera tion for knowledge-intensive NLP ta sks. Advances in Neural Information Processing Systems, 33, 9459–9474. https://proceedings.neurips.c/pa per/2020/ha sh/6b493230205f780e1bc26945df7481e5-Abstra ct.html Pfeifer, M., & Ma rohl, V. P. (2023). Centra lBa nkRoBERTa : A fine-tuned la rge la ngua ge model for centra l bank communica tions. The Journal of Finance and Data Science, 9, 100114. https://doi.org/10.1016/j.jfds.2023.100114 Silva , T. C., Moriya , K., & Veyrune, R. M. (2025). From text to quantified insights: A large-scale LLM analysis of central bank communication (IMF Working Pa per No. 2025/109). Internationa l Moneta ry Fund. https://doi.org/10.5089/9798229013802.001 12 Appendix A. Implementation Audit Audit item Observed implementation Assessment Date coverage 80 matching dates after the August 2016 filter Internally aligned study sample Score denominator Hawk + dove + neutral; out excluded Consistent with published iSent-style denominator Intensity calibration Per-document signal means; global means saved separately Article reports exact live logic Decision-only rule Prompt instructs neutral, but some historical outputs violate it Requires post-processing or gold-set audit Historical anchors Three heuristic dates in the prompt Guides LLM but does not constrain final score Structural score Direction x explicitness Correctly separated from tone FAISS Index builder exists behind optional flag Not queried in the classification path Incremental cache Only unseen dates are classified Efficient, but requires model-version metadata Raw corpus/config Expected by code but absent from supplied archive Blocks exact end-to-end replication Feature CSV 78 rows through April 2026 JSON adds June and August 2026, yielding 80 study dates Appendix Table A1. Implementation audit of the uploaded Copom sentiment project. Appendix B. Scoring Interpretation The index is a tone mea sure, not a probability. A score of +0.23 does not mean a 23% probability of a ra te increa se. It means that the weighted ba lance of ha wkish and dovish sentence counts, after dilution by neutra l content, equa ls +0.23 under the current a ggrega tion rule. L ikewise, zero ca n a rise from genuinely ba la nced langua ge, offsetting we ighted counts, or a sta tement domina ted by neutra l segments. Cross-period compa risons should consider document length a nd source forma tting. Ea rlie r sta tements within the reta ined sample can be shorter tha n recent sta tements, ma king their scores more discrete and more sensitive to a single cla ssification. The score is therefore most relia ble a s one input to a broader communication da shboa rd that a lso reports sentence counts, guida nce, uncerta inty, a nd the underlying excerpts.