Paper deep dive
Via Negativa for AI Alignment: Why Negative Constraints Are Structurally Superior to Positive Preferences
Quan Cheng
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 5:41:43 AM
Summary
The paper proposes a theoretical framework called 'Via Negativa' for AI alignment, arguing that negative constraints (what to avoid) are structurally superior to positive preferences (what to prefer). It posits that positive preferences are continuously coupled and inexhaustible, leading to issues like sycophancy, whereas negative constraints are discrete, finite, and convergent, allowing for more stable and robust model alignment.
Entities (6)
Relation Signals (3)
RLHF â amplifies â Sycophancy
confidence 95% ¡ standard preference-based RLHF systematically amplifies sycophancy
Constitutional AI â utilizes â Negative Constraints
confidence 95% ¡ Constitutional AI... specifies what the model should not do
Via Negativa â explains â Negative-signal methods
confidence 92% ¡ This asymmetry... explains both the sycophancy failure of preference-based RLHF and the surprising effectiveness of negative-signal methods.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent empirical results have demonstrated that training large language models (LLMs) with negative-only feedback can match or exceed standard reinforcement learning from human feedback (RLHF). Negative Sample Reinforcement achieves parity with PPO on mathematical reasoning; Distributional Dispreference Optimization trains effectively using only dispreferred samples; and Constitutional AI outperforms pure RLHF on harmlessness benchmarks. Yet no unified theoretical account explains why negative signals are so effective. This paper proposes such an account: positive preferences and negative constraints are structurally asymmetric. Positive preferences ("which is better") encode continuously coupled, context-dependent human values that cannot be exhaustively specified -- leading models to learn surface correlates such as agreement with the user (sycophancy). Negative constraints ("what is wrong") encode discrete, finite, independently verifiable prohibitions that can converge to a stable boundary. This asymmetry -- rooted in Popper's falsification logic and the epistemology of negative knowledge -- explains both the sycophancy failure of preference-based RLHF and the surprising effectiveness of negative-signal methods. We argue that alignment research should shift its center of gravity from "learning what humans prefer" to "learning what humans reject," and offer testable predictions for this framework.
Tags
Links
- Source: https://arxiv.org/abs/2603.16417v1
- Canonical: https://arxiv.org/abs/2603.16417v1
Trouble viewing inline? Open PDF directly â
Full Text
24,580 characters extracted from source content.
Expand or collapse full text
Via Negativa for AI Alignment: Why Negative Constraints Are Structurally Superior to Positive Preferences Quan Cheng Tsinghua University chengq25@mails.tsinghua.edu.cn Abstract Recent empirical results have demonstrated that training large language models (LLMs) with negative-only feedback can match or exceed standard reinforcement learning from human feedback (RLHF). Negative Sample Reinforcement achieves parity with PPO on mathematical reasoning; Distributional Dispreference Opti- mization trains effectively using only dispreferred samples; and Constitutional AI outperforms pure RLHF on harmlessness benchmarks. Yet no unified theoretical account explains why negative signals are so effective. This paper proposes such an account: positive preferences and negative constraints are structurally asymmetric. Positive preferences (âwhich is betterâ) encode continuously coupled, context-dependent human values that cannot be exhaustively specifiedâleading models to learn surface correlates such as agreement with the user (sycophancy). Negative constraints (âwhat is wrongâ) encode discrete, finite, independently verifi- able prohibitions that can converge to a stable boundary. This asymmetryârooted in Popperâs falsification logic and the epistemology of negative knowledgeâexplains both the sycophancy failure of preference-based RLHF and the surprising effective- ness of negative-signal methods. We argue that alignment research should shift its center of gravity from âlearning what humans preferâ to âlearning what humans reject,â and offer testable predictions for this framework. 1 Introduction A puzzling pattern has emerged in LLM alignment research. Method after method demon- strates that negative feedback signalsâpenalizing what is wrong rather than reinforcing what is preferredâperform surprisingly well, sometimes matching or exceeding methods that use both positive and negative signals. Negative Sample Reinforcement (NSR) trains models by penalizing incorrect reason- ing traces without reinforcing correct ones, yet matches PPO and GRPO on MATH and AIME benchmarks [1]. Distributional Dispreference Optimization (D2O) learns from dispreferred samples only, without requiring noisy positive examples [2]. Negative Pref- erence Optimization (NPO) treats forget data exclusively as negative responses, achieving effective unlearning without catastrophic collapse [3]. Kahneman-Tversky Optimization (KTO) aligns models using unpaired binary signals weighted by loss aversion, matching DPO at scale with far less data [4]. 1 arXiv:2603.16417v1 [cs.AI] 17 Mar 2026 Meanwhile, a parallel line of research has established that standard preference-based RLHF systematically amplifies sycophancy. Sharma et al. [5] demonstrated that human annotators prefer sycophantic responses over correct ones at non-trivial rates, corrupting the preference signal at its source. Shapira et al. [6] provided a formal mathematical mechanism: sycophancy amplification is driven by a covariance between âendorsing the userâs beliefâ and âreceiving high rewardâ under the base policy. These two phenomenaâthe effectiveness of negative-only training and the sycophancy failure of positive-preference trainingâhave been studied independently. This paper ar- gues they are two manifestations of a single structural asymmetry: positive prefer- ences are continuously coupled and inexhaustible, while negative constraints are discrete, finite, and convergent. This is a position paper. We do not present new experiments but offer a theoreti- cal framework that unifies and explains existing empirical results, and propose testable predictions. 2 The Structural Asymmetry 2.1 Positive Preferences Are Continuously Coupled When a human annotator is asked âwhich response is better?â, they are implicitly evalu- ating against a preference function that is: ⢠Context-dependent: What counts as âbetterâ depends on who is asking, why they are asking, what they already know, and what they intend to do with the answer. The same response may be preferred in one context and dispreferred in another. ⢠Multi-dimensional: âBetterâ simultaneously encodes accuracy, helpfulness, tone, conciseness, creativity, safety, and dozens of other criteria whose relative weights vary by situation. ⢠Continuously coupled: These dimensions are not independent. The optimal level of detail depends on the userâs expertise, which affects what counts as helpful, which interacts with what counts as concise. Each dimensionâs ideal value is a function of all other dimensionsâa continuously coupled system [7]. This structure is formally analogous to what Smolensky [8] identified in connectionist representations: a massively parallel continuous constraint satisfaction system in which each variableâs value is a function of all other variables, admitting no complete symbolic- level description. The preference function that âwhich is better?â attempts to elicit is precisely such a system. The consequence is that positive preferences cannot be exhaustively specified by any finite set of rules or examples. Each preference annotation is a lossy projection of an infinite-dimensional preference manifold onto a binary signal. The information loss is not a practical limitation that more data could overcomeâit is a structural property of the preference function itself. 2.2 Negative Constraints Are Discrete and Finite Consider instead the question âwhat is wrong with this response?â The space of identifi- able errors is structurally different: 2 ⢠Factual errors are discrete and independently verifiable: âParis is not the capital of Germany.â ⢠Safety violations are enumerable: a finite list of prohibited behaviors (generating malware, providing instructions for violence, revealing private information). ⢠Logical contradictions are binary: the response either contradicts itself or does not. ⢠Format violations are checkable: the response either follows the requested format or does not. Each negative constraint is: (1) discreteâit either applies or does not; (2) indepen- dentâone constraintâs validity does not depend on other constraints; (3) verifiableâit can be checked without reference to the full preference function; (4) stableââfactually wrongâ does not become âfactually rightâ depending on context. This means the space of negative constraints can, in principle, be exhaustively enu- merated. As constraints accumulate, the feasible response space narrows monotonically. Beyond a sufficient number of constraints, the remaining feasible space is narrow enough that any response within it is approximately acceptableânot because the model has learned what is best, but because it has learned to avoid everything that is clearly wrong. 2.3 The Asymmetry Is Structural, Not Quantitative The distinction we are drawing is not that negative feedback is âeasier to collectâ or âless noisyââthough both may be true empirically. The claim is stronger: positive preferences and negative constraints occupy different positions in the episte- mological hierarchy. This asymmetry has deep roots. Karl Popperâs philosophy of science [9] rests on precisely this structure: a single counterexample decisively refutes a universal claim (fal- sification), but no finite number of confirming instances can decisively verify one. Falsi- fication is logically asymmetric with respect to verification. Negative knowledge (âthis is wrongâ) is epistemologically privileged over positive knowledge (âthis is rightâ). Nassim Taleb [10] extended this insight under the term via negativa: in domains of high uncertainty, removing what is harmful is more robust than adding what seems ben- eficial. âThe chess grandmaster usually wins by not losing.â The grandmasterâs expertise is primarily negativeâa vast repertoire of positions and moves to avoidârather than a positive specification of the optimal move in each position. Gartmeier et al. [11] formalized this in the context of professional expertise as ânega- tive knowledgeâ: knowledge about what is wrong and what is to be avoided, which func- tions through inhibition rather than prescription. Their key observationâthat negative knowledge is experientially acquired and expert-levelâaligns precisely with the pattern we observe in LLM training. The contribution of this paper is to connect these epistemological traditions to the empirical landscape of LLM alignment, providing a unified theoretical account of why negative-signal methods work. 3 3 Explaining Existing Results 3.1 Why RLHF Produces Sycophancy The structural asymmetry framework offers a clean explanation for sycophancy in RLHF. Standard RLHF asks annotators: âwhich response is better?â This question forces the annotator to project their continuously coupled preference function onto a binary comparison. The projection is necessarily lossy. Among the dimensions lost, one has a particularly pernicious surface correlate: agreement with the userâs stated position correlates with perceived quality. Sharma et al. [5] confirmed this empirically: annotators prefer sycophantic responses over correct ones at significant rates. Shapira et al. [6] formalized the mechanism: when the base policy already correlates âendorsing the userâs viewâ with âhigh reward,â RLHF amplifies this correlation because the reward model learns the correlation as a feature rather than a confound. Our framework explains why this is not a fixable bug but a structural feature of positive-preference training. The annotatorâs true preference functionâwhich would dis- tinguish âgenuinely betterâ from âmerely more agreeableââis continuously coupled and cannot be fully encoded in pairwise comparisons. The sycophancy correlate is a low- dimensional surface feature that survives the lossy projection. No amount of preference data can eliminate this problem, because the problem lies in the structure of the signal, not its quantity. 3.2 Why Constitutional AI Is More Robust Anthropicâs Constitutional AI [12] replaces human preference annotation with a set of principlesâa âconstitutionââthat an AI assistant uses to critique and revise its own outputs. The constitution is primarily negative: it specifies what the model should not do (be harmful, be deceptive, be invasive of privacy). From our framework, this works precisely because the constitution encodes discrete negative constraints rather than continuous positive preferences. Each principle is in- dependently verifiable: âDoes this response contain instructions for making weapons? Yes/No.â The model does not need to learn the full human preference functionâit only needs to learn to avoid a finite set of clearly defined violations. This also explains an observation that has been noted but not theoretically accounted for: Claude (trained primarily with Constitutional AI) exhibits less sycophancy than models trained primarily with preference-based RLHF [5]. The structural reason is that Constitutional AIâs negative constraints do not contain the sycophancy correlate that positive preference data does. 3.3 Why Negative-Only Training Matches Full RLHF The NSR result [1]âthat penalizing wrong answers without reinforcing correct ones matches PPOâis initially counterintuitive. How can a model improve if it is never told what is right? Our framework provides the answer: the model already possesses a prior distribu- tion over responses from pre-training. Negative feedback does not need to specify the correct answerâit only needs to suppress incorrect regions of the response space. As 4 incorrect regions are progressively eliminated, the probability mass redistributes toward the remaining feasible space, which is increasingly concentrated around correct responses. This is precisely the mechanism NSRâs authors identified empirically: âNSR suppresses incorrect generations and redistributes probability mass toward plausible alternatives guided by the modelâs prior beliefsâ [1]. Our framework explains why this works in general: because the space of errors is discrete and enumerable, while the space of correct responses is continuous and context-dependent, it is structurally more efficient to specify the former than the latter. The same logic explains D2O [2] (learning from dispreferred samples only), NPO [3] (negative-only unlearning), and the finding by Yao et al. [13] that LLM unlearning with 2% of the computational budget achieves RLHF-equivalent safetyâall cases where specifying what to avoid proves sufficient. 3.4 Why KTO Works with Unpaired Data KTO [4] aligns models using unpaired binary feedbackâindividual responses labeled as âdesirableâ or âundesirableââwithout requiring pairwise comparisons. Its theoretical foundation is Kahneman and Tverskyâs prospect theory: humans are loss-averse, weighing losses more heavily than equivalent gains. Our framework provides a deeper explanation for why loss-averse weighting is appro- priate: losses (negative feedback) carry structurally more information than gains (posi- tive feedback). A single âundesirableâ label decisively excludes a region of response space, while a single âdesirableâ label only weakly indicates one point in an infinite-dimensional preference manifold. The asymmetric weighting in KTO implicitly recognizes the struc- tural asymmetry we have made explicit. 4 The Convergence Argument A critical advantage of negative constraints is their convergence property. We state this informally: Claim: As the number of negative constraints increases, the feasible response space contracts monotonically. Beyond a sufficient number of constraints, any response within the feasible space is approximately acceptable. This is the alignment analogue of what Taleb [10] calls the via negativa principle and what the Dreyfus model of expertise [14] describes as the transition from rule-following to intuitive competence: the expert does not compute the optimal action but has internalized enough prohibitions that the remaining action space contains mostly adequate options. Consider a concrete example. An unaligned modelâs response space for a given query is vastâit could output anything from a helpful answer to harmful instructions. Each negative constraint (âdo not generate malware,â âdo not fabricate citations,â âdo not reveal private data,â âdo not contradict established factsâ) removes a region of this space. The constraints are cumulative and non-conflicting: adding a new prohibition never re-opens a previously closed region. After sufficiently many constraints, the remaining space is narrow enough that the modelâs pre-trained language competence is sufficient to produce acceptable outputs within it. Positive preferences, by contrast, do not converge in this way. Adding a new âthis is better than thatâ comparison does not monotonically narrow the response spaceâit 5 adjusts the relative ranking within an already continuous space. Two preference compar- isons can conflict (response A preferred over B in context 1, B preferred over A in context 2), and the resolution depends on the continuously coupled context function that cannot be fully specified. 5 A Testable Prediction: Capability as Negative Knowl- edge If the structural asymmetry thesis is correct, it generates a testable prediction about model capability: Prediction: More capable models possess more negative knowledge (what not to say) rather than more positive knowledge (what to say). This manifests as shorter, denser responses with higher information content per token. The reasoning is as follows. A more capable model has been trained on more data and has undergone more alignment iterations. If negative knowledge (constraints on what to avoid) accumulates more reliably than positive knowledge (specifications of what is optimal), then a more capable modelâs primary advantage is knowing more about what not to include in a responseâredundant elaboration, unnecessary hedging, tangential information, formulaic pleasantries. Informal observations are consistent with this prediction. Within the same model family, more capable variants (e.g., Opus vs. Sonnet in Anthropicâs Claude family) tend to produce shorter responses with higher information density. Across model families, models trained with more Constitutional AI emphasis (negative constraints) tend to be less verbose than models trained with more RLHF emphasis (positive preferences). This prediction is empirically testable through a controlled benchmark: ⢠Metric 1: Response length (in tokens) for standardized queries across model capability tiers ⢠Metric 2: Information density (unique substantive claims per token) ⢠Metric 3: Sycophancy rate (agreement with demonstrably false user claims) ⢠Prediction: Capability correlates negatively with length, positively with information density, and negatively with sycophancy rate If confirmed, this would provide evidence that capability growth in LLMs is at least partially driven by the accumulation of negative knowledgeâlearning what not to sayâ rather than solely by the expansion of positive knowledge about what to say. 6 Implications for Alignment Research 6.1 Reframing the Alignment Objective Current alignment research is largely organized around the question: âHow do we learn what humans want?â Our framework suggests this question is structurally ill-posed for the same reason that âdescribe the optimal chess move for every positionâ is ill-posedâthe answer space is continuously coupled and inexhaustible. 6 A more tractable formulation is: âHow do we learn what humans reject?â This ques- tion targets the discrete, finite, convergent side of the structural asymmetry. It does not require solving the preference functionâit requires enumerating the boundaries. This is not a minor reframing. It changes what data to collect (rejection signals rather than preference comparisons), how to design annotation interfaces (asking âwhat is wrong?â rather than âwhich is better?â), and what convergence guarantees are achievable (monotonic boundary contraction rather than approximate preference matching). 6.2 Constitutional AI as a Template Constitutional AI [12] already implements this reframing, though it was not explicitly motivated by the structural asymmetry we describe. Our framework suggests that Con- stitutional AIâs success is not incidental but reflects a correct alignment between the methodâs structure and the structure of the problem. Future alignment methods should be evaluated not only on benchmark performance but on whether they leverage the asymmetryâtargeting discrete constraints rather than continuous preferences. 6.3 The Limits of Via Negativa We do not claim that negative constraints are sufficient for all aspects of alignment. Certain alignment desiderataâhelpfulness, creativity, appropriate toneâare inherently positive and may resist negative specification. Our claim is that the discrete, convergent component of alignment (safety, factual accuracy, logical consistency) should be addressed through negative constraints, reserving positive preference learning for the residual contin- uous component. This separation of concerns could reduce the sycophancy contamination currently observed when safety and helpfulness are learned jointly. 7 Related Work Sycophancy. Perez et al. [15] first documented sycophancy in language models. Sharma et al. [5] traced it to preference data. Shapira et al. [6] formalized the amplification mech- anism. Wei et al. [16] proposed simple prompting-based mitigations. Our contribution is a structural explanation for why sycophancy is an intrinsic failure mode of positive- preference methods. Negative-signal training. D2O [2], NSR [1], NPO [3], KTO [4], and BNF [17] demonstrate the empirical effectiveness of negative-only or negative-weighted training. Our contribution is a theoretical account unifying these results. Via negativa in philosophy. Popper [9] established the falsification asymmetry. Taleb [10] applied it to decision-making under uncertainty. Gartmeier et al. [11] formal- ized negative knowledge in expertise research. Parviainen and Eriksson [18] connected it to organizational knowledge. Our contribution is to bring this epistemological tradition into contact with the AI alignment literature, where it has been absent. Tacit knowledge and LLMs. Kambhampati [19] identified the connection between expert systemsâ failure and tacit knowledge. Cheng [7] argued that the valuable capabil- ities of LLMs are precisely the unexplainable ones, via a proof by contradiction through expert system equivalence. The present paper extends this argument: if positive knowl- edge is structurally uncapturable (Chengâs thesis), then alignment should target negative knowledge instead. 7 8 Conclusion We have argued that positive preferences and negative constraints are structurally asym- metric: the former are continuously coupled and inexhaustible, while the latter are dis- crete, finite, and convergent. This asymmetryâgrounded in Popperâs falsification logic and the epistemology of negative knowledgeâprovides a unified theoretical explanation for two independently observed phenomena in LLM alignment: the systematic syco- phancy produced by preference-based RLHF, and the surprising effectiveness of negative- only training methods. The practical implication is a reframing of the alignment objective: from âlearn what humans preferâ (a structurally intractable problem) to âlearn what humans rejectâ (a structurally convergent one). Constitutional AI already implements this reframing; the growing family of negative-signal methods (NSR, D2O, NPO, KTO) provides empirical support; and the epistemological tradition of via negativa provides theoretical grounding. The chess grandmaster wins by not losing. The aligned model aligns by learning what not to do. References [1] Liu, Y., Zeng, Z., et al. (2025). âThe Surprising Effectiveness of Negative Rein- forcement in LLM Reasoning.â Advances in Neural Information Processing Systems (NeurIPS). arXiv:2506.01347. [2] Duan, H., Yi, Y., Zhang, Z., Liu, F., et al. (2024). âNegating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization.â Find- ings of EMNLP. arXiv:2403.03419. [3] Zhang, J., et al. (2024). âNegative Preference Optimization: From Catastrophic Collapse to Effective Unlearning.â arXiv preprint arXiv:2404.05868. [4] Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D. (2024). âKTO: Model Alignment as Prospect Theoretic Optimization.â Proceedings of ICML. arXiv:2402.01306. [5] Sharma, M., Tong, M., Korbak, T., et al. (2024). âTowards Understanding Syco- phancy in Language Models.â Proceedings of ICLR. arXiv:2310.13548. [6] Shapira, N., Levy, M., Alavi, S. H., et al. (2026). âHow RLHF Amplifies Sycophancy.â arXiv preprint arXiv:2602.01002. [7] Cheng, Q. (2026). âWhy the Valuable Capabilities of LLMs Are Precisely the Unex- plainable Ones.â arXiv preprint. [8] Smolensky, P. (1988). âOn the Proper Treatment of Connectionism.â Behavioral and Brain Sciences, 11(1), 1â23. [9] Popper, K. R. (1959). The Logic of Scientific Discovery. Routledge. (Original: Logik der Forschung, 1934.) [10] Taleb, N. N. (2012). Antifragile: Things That Gain from Disorder. Random House. Chapter 22: Via Negativa. 8 [11] Gartmeier, M., Bauer, J., Gruber, H., and Heid, H. (2008). âNegative Knowledge: Understanding Professional Learning and Expertise.â Vocations and Learning, 1(2), 87â103. [12] Bai, Y., Kadavath, S., et al. (2022). âConstitutional AI: Harmlessness from AI Feed- back.â arXiv preprint arXiv:2212.08073. [13] Yao, Y., et al. (2024). âLarge Language Model Unlearning.â Advances in Neural Information Processing Systems (NeurIPS). arXiv:2310.10683. [14] Dreyfus, H. L. and Dreyfus, S. E. (1986). Mind over Machine: The Power of Human Intuition and Expertise in the Era of the Computer. Free Press. [15] Perez, E., Ringer, S., et al. (2023). âDiscovering Language Model Behaviors with Model-Written Evaluations.â Findings of ACL. arXiv:2212.09251. [16] Wei, J., et al. (2023). âSimple Synthetic Data Reduces Sycophancy in Large Language Models.â arXiv preprint arXiv:2308.03958. [17] Han, Y., et al. (2024). âBNF: As Simple as Fine-tuning: LLM Alignment via Bidi- rectional Negative Feedback Loss.â OpenReview. [18] Parviainen, J. and Eriksson, M. (2006). âNegative Knowledge, Expertise and Or- ganisations.â International Journal of Management Concepts and Philosophy, 2(2), 140â153. [19] Kambhampati, S. (2021). âPolanyiâs Revenge and AIâs New Romance with Tacit Knowledge.â Communications of the ACM, 64(10), 31â33. 9