Paper deep dive
Baby Scale: Investigating Models Trained on Individual Children's Language Input
Steven Y. Feng, Alvin W. M. Tan, Michael C. Frank
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/1/2026, 1:35:43 AM
Summary
This paper investigates the data efficiency of language models (LMs) by training them on child-scale datasets from the BabyView corpus. The authors analyze how model performance scales with input size, the impact of linguistic features on model quality, and the correlation between model likelihoods and human child language acquisition outcomes. Results indicate that while grammatical knowledge scales well at child-scale regimes, semantic and world knowledge tasks show lower scaling, and performance is significantly influenced by distributional and interactional linguistic features.
Entities (6)
Relation Signals (3)
GPT-2 ā evaluatedon ā Zorro
confidence 95% Ā· We evaluate several different training conditions... Zorro assesses the grammatical capabilities
BabyView ā usedtotrain ā GPT-2
confidence 95% Ā· We train two main model families... GPT-2... using transcripts from the BabyView dataset
TinyDialogues ā usedtoinvestigate ā Scaling Properties
confidence 90% Ā· to investigate further scaling properties, we also fit models on varying subsets of... the TinyDialogues corpus
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Modern language models (LMs) must be trained on many orders of magnitude more words of training data than human children receive before they begin to produce useful behavior. Assessing the nature and origins of this "data gap" requires benchmarking LMs on human-scale datasets to understand how linguistic knowledge emerges from children's natural training data. Using transcripts from the BabyView dataset (videos from children ages 6-36 months), we investigate (1) scaling performance at child-scale data regimes, (2) variability in model performance across datasets from different children's experiences and linguistic predictors of dataset quality, and (3) relationships between model and child language learning outcomes. LMs trained on child data show acceptable scaling for grammar tasks, but lower scaling on semantic and world knowledge tasks than models trained on synthetic data; we also observe substantial variability on data from different children. Beyond dataset size, performance is most associated with a combination of distributional and interactional linguistic features, broadly consistent with what makes high-quality input for child language development. Finally, model likelihoods for individual words correlate with children's learning of those words, suggesting that properties of child-directed input may influence both model learning and human language development. Overall, understanding what properties make language data efficient for learning can enable more powerful small-scale language models while also shedding light on human language acquisition.
Tags
Links
- Source: https://arxiv.org/abs/2603.29522v1
- Canonical: https://arxiv.org/abs/2603.29522v1
Trouble viewing inline? Open PDF directly ā
Full Text
103,441 characters extracted from source content.
Expand or collapse full text
Baby Scale: Investigating Models Trained on Individual Childrenās Language Input Steven Y. Feng, Alvin W.M. Tan, Michael C. Frank Stanford University syfeng,tanawm,mcfrank@stanford.edu Abstract Modern language models (LMs) must be trained on many orders of magnitude more words of training data than human children receive before they begin to produce useful behavior. Assessing the nature and origins of this ādata gapā requires benchmarking LMs on human-scale datasets to understand how linguistic knowledge emerges from childrenās natural training data. Using transcripts from the BabyView dataset (videos from children ages 6ā36 months), we investigate (1) scaling performance at child-scale data regimes, (2) variability in model performance across datasets from different childrenās experiences and linguistic predictors of dataset quality, and (3) relationships between model and child language learning outcomes. LMs trained on child data show acceptable scaling for grammar tasks, but lower scaling on semantic and world knowledge tasks than models trained on synthetic data; we also observe substantial variability on data from different children. Beyond dataset size, performance is most associated with a combination of distributional and interactional linguistic features, broadly consistent with what makes high-quality input for child language development. Finally, model likelihoods for individual words correlate with childrenās learning of those words, suggesting that properties of child-directed input may influence both model learning and human language development. Overall, understanding what properties make language data efficient for learning can enable more powerful small-scale language models while also shedding light on human language acquisition.ā Code and data: https://github.com/styfeng/babyscale-LM 1 Introduction Large language models (LLMs) achieve impressive capabilities by training on enormous datasets consisting of billions or trillions of tokens. In contrast, human children acquire language with orders of magnitude less linguistic input. Children receive between 10710^7 and 10810^8 words during early development ā far smaller than the datasets used to train modern LLMs, which are often 101210^12ā101310^13 words (Frank, 2023a; Warstadt et al., 2023). This ādata gapā likely results from many causes, including language model architecture, the nature of their training, and the nature of the data they are trained on. Understanding what makes children more efficient learners ā and potentially importing these insights to language models ā requires the evaluation of language models on datasets from human language acquisition. Small language models (SLMs) trained on human-scale datasets have emerged as an important tool for investigating data efficiency. For example, the BabyLM challenges constrain training data to developmentally plausible budgets and ask participants to develop architectures and training methodologies that learn efficiently under such conditions (Warstadt et al., 2023; Hu et al., 2024). These competitions have led to substantial advances (e.g., Charpentier and Samuel, 2024; Martinez et al., 2023). Yet, by holding the training data constant ā a necessary feature of these challenges ā they prevent researchers from answering an important but distinct set of questions about what makes some language datasets more efficient for learning than others, and to what extent efficient learning might be driven by certain types of data. Here, we use data from childrenās learning environments to study what makes data āgoodā for learning at human scales. We are inspired by developmental psychology studies of the importance of dataset quantity vs. quality in childrenās language learning (Anderson et al., 2021; Masek et al., 2021; Rowe, 2012; Rowe and Snow, 2020), which suggest that children learn more from input that contains complex sentence structures and diverse vocabulary. Data are a key bottleneck in studying child language. Although several tens of millions of words of data from childrenās naturalistic conversations are available in the CHILDES archive (MacWhinney, 2014), it is challenging to compare across them as they are collected from very different studies with different methods. Further, these corpora typically do not have outcome measures on language acquisition in the individual children who contribute, meaning that the datasets cannot be linked to individual childrenās learning progress. We address these issues by using the BabyView dataset (Long et al., 2024), a collection of egocentric recordings (with transcripts) capturing naturalistic language experiences by children across multiple families. Because each family represents a distinct linguistic environment and context for language learning, BabyView enables us to train SLMs on individual familiesā data to simulate āsynthetic learnersā exposed to different language learning environments (following Wang et al., 2023; Qin et al., 2024). Our setting allows us to systematically examine several key research questions: (1) How does language model performance scale with the amount of input received by a single child, or a mixture of children? (2) How much variability arises across different familiesā language environments, and which linguistic features of the input data predict language model performance? (3) Do properties of child-directed input relate to childrenās language outcomes? Our results reveal that scaling performance varies across tasks: while scaling remains visible in grammar even at child-scale data regimes, semantic and world knowledge tasks scale much less well. Further, variability across families is substantial and explained by factors beyond dataset size, including distributional and interactional properties. Finally, model likelihoods relate to individual childrenās outcomes. Together, these results provide new insight into the role of input data in language learning and highlight the potential of child-scale datasets as a lens for understanding human and machine language acquisition. 2 Background and Related Work Scaling laws have become a central framework for understanding LM performance. Prior work demonstrates that performance often follows power-law relationships with model size, dataset size, and compute. However, most empirical studies examine regimes involving billions of tokens (Kaplan et al., 2020; Henighan et al., 2020; Hoffmann et al., 2022; Muennighoff et al., 2023). Further, there is a body of work on optimal data mixing and selection (Tirumala et al., 2023; Gao et al., 2020; Longpre et al., 2023; Feng et al., 2024a), but also at much larger scales. While there are studies examining how scaling laws vary based on the data being provided (Sorscher et al., 2022; Pandey, 2024; Brill, 2024), much less is known about scaling behavior in extremely small data regimes. Understanding whether similar scaling dynamics emerge at child-scale data sizes ā and how they are modulated by aspects of dataset quality ā can provide insight into both model efficiency and the role of input data. Much recent work on the properties of SLMs comes from the BabyLM challenge (Warstadt et al., 2023; Hu et al., 2024), which investigates SLM training using fixed budgets of 10M and 100M tokens in a heterogeneous mix of child-directed speech with other text data. Other work has used this kind of paradigm to understand questions related to dataset āqualityā (Zhang et al., 2021; Huebner et al., 2021) ā often by using a ācontrolled rearingā design in which models are held constant while varying datasets (Frank, 2023b). These designs allow inferences about modelsā inductive biases (Kallini et al., 2024), systematic asymmetries (Hu et al., 2025), their ability to generalize to unseen grammatical structures (Misra and Mahowald, 2024), and to learn from multilingual input (Constantinescu et al., 2025), among other topics. Our work uses this design to investigate the informational value of child-directed speech. Initial evidence suggested that a pure corpus of child-directed speech might be less informative than the BabyLM mixture, but this study did not provide evidence for why (Feng et al., 2024b). Qin et al. (2024) took important steps by training a range of SLMs on data from three children. However, the datasets being compared were from very different studies and corpora, and no attempt was made to interpret between-child differences. In fact, a rich literature on childrenās language development suggests that there are substantial differences in language acquisition between children (Frank et al., 2021; Kidd and Donnelly, 2020), and that these differences are related to differences in childrenās language input (Hart et al., 1997; Weisleder and Fernald, 2013; Coffey and Snedeker, 2025). While initial evidence on this question focused on differences in the quantity of speech that children hear across households (Hart et al., 1997), subsequent investigations have also emphasized that other measures of linguistic quality ā including lexical and syntactic diversity ā may be even more important (Masek et al., 2021; Anderson et al., 2021; Rowe, 2012; Rowe and Snow, 2020). In sum, our work here builds on prior work investigating ācontrolled rearingā in SLMs to understand aspects of dataset quality at child-scale. Inspired by research on child language acquisition, and using dense longitudinal transcripts from a single study, we analyze the features that relate to the effectiveness of child-directed speech as training data. 3 Methods 3.1 Datasets Our training data consists of conversation between children and caregivers, set up so that each line corresponds to a single conversation. Each training dataset was split into 85/15 train/val splits. Examples and preprocessing details can be found in Appendix A. BabyView. The BabyView dataset consists of egocentric video recordings capturing childrenās everyday experiences at home (Long et al., 2024). The 2025.2 release of the dataset includes ā¼1500 1500 hours of video, speech transcripts using WhisperX large-v3 (Radford et al., 2023), and speaker role diarization using VTC 2.0 (Kunze et al., 2025). For this study, we use data only from 20 monolingual English-speaking families (BāVEānāgBV_Eng), identified through a combination of automatic and manual filtering. These individual family transcripts range from ā¼ 1000 to ā¼ 725k tokens, with an aggregate total of ā¼ 2.8M tokens in BāVEānāgBV_Eng. TinyDialogues. The total amount of data in BāVEānāgBV_Eng is very limited, but this limitation reflects the overall scarcity of child-directed speech transcripts; the CHILDES archive contains only ā¼ 29M English words (Huebner et al., 2021; MacWhinney, 2014). Thus, to investigate further scaling properties, we also fit models on varying subsets of a larger dataset (up to 200M words) of synthetic child-directed dialogues, the TinyDialogues corpus (Feng et al., 2024b). Inspired by the TinyStories project (Eldan and Li, 2023), these dialogues are generated by GPT-4 to feature diverse situations and vocabulary in short, simple dialogues between children (of varying ages) and caregivers. 3.2 Experimental Conditions We evaluate several different training conditions, totaling 25 experiments: ⢠Individual-family models: one model trained per individual BāVEānāgBV_Eng family (20). ⢠Pooled mixtures: models trained on size-controlled mixtures of differing numbers of BāVEānāgBV_Eng families (2, 5, 10, and all 20), matching the total linguistic data (ā¼ 725k tokens) of the largest individual BāVEānāgBV_Eng family. We proportionally sample conversations from the individual families to most closely reach the desired token count. ⢠All-families model: trained on the entirety (all families) of BāVEānāgBV_Eng. These conditions allow us to compare the effects of child-scale dataset size and diversity on language model learning. 3.3 Model Architectures We train two main model families. First, we train the autoregressive LM, GPT-2 (Radford et al., 2019), with 124M parameters (small version), following prior ācontrolled rearingā work (Kallini et al., 2024; Misra and Mahowald, 2024; Qin et al., 2024). Second, we train the hybrid LM, GPT-BERT, which uses masked next-token prediction (MNTP) combining autoregressive and masked objectives within a single Transformer architecture (Charpentier and Samuel, 2024); GPT-BERT was found to be effective for small-scale training on the BabyLM challenge. Training details (including hyperparameters) can be found in Appendix B. We investigated two sizes of each model, to investigate whether performance might vary at such small data scales, and in case of overfitting. We looked at GPT-2 small (124M), GPT-2 mini (39M), GPT-BERT base (119M), and GPT-BERT small (30M). For both model architectures, we train a separate tokenizer on the entire BāVEānāgBV_Eng data.111We do not train separate tokenizers for every individual experiment, as many individual BāVEānāgBV_Eng families contained too few unique types. As such, filtering our evaluation data using the intersection of the individual familiesā vocabulary would eliminate the vast majority of evaluation examples. We train five separate seeds of each model for every experiment, reporting averaged results and seed variability. 3.4 Evaluation Benchmarks Drawing from the second BabyLM competition evaluation suite (Hu et al., 2024) and prior related work such as Feng et al. (2024b), we selected benchmarks suitable to the scale and domain of our dataset, discussed below. Note that we filtered the evaluation data by our BabyView datasetās vocabulary. In particular, an evaluation example was kept only if every token in the example appeared at least once in BāVEānāgBV_Eng. Zorro assesses the grammatical capabilities of SLMs via forced-choice judgments between minimal pairs of grammatical and ungrammatical sentences created from a limited vocabulary (Huebner et al., 2021). We report average accuracy across individual Zorro tasks. Word Similarity measures the ability of models to capture semantic similarities between pairs of words, allowing us to assess the semantic knowledge of our models (Zhuang et al., 2023). We report Spearman correlations between human and model similarity judgments. More details in Appendix C. COMPS measures if models can demonstrate property inheritance, and infer that properties of superordinate concepts are inherited by subordinate concepts represented by nonce words (Misra et al., 2023). Similar to Zorro, models must choose between minimal pairs of sentences, and we report accuracy of assigning higher probability to the correct sentence. EWoK evaluates modelsā basic world knowledge by testing whether they can distinguish plausible from implausible scenarios given a context (Ivanova et al., 2025). The benchmark spans multiple domains of everyday knowledge, including social, physical, and spatial reasoning. We report accuracy of assigning higher probability to the correct continuation. 3.5 Linguistic Feature Analysis To better understand why some BabyView family datasets produce stronger models than others, we compute a variety of linguistic features for each dataset (a full catalog of all 175 linguistic features is provided in Table LABEL:app_tab:ling_feature_catalog_detailed in Appendix F). These features cover several categories of potential predictors of model performance: ⢠Lexical/distributional scale: e.g., token and conversation counts, TTR/MATTR, entropy, Zipf/skewness, and related concentration statistics. ⢠Syntactic composition: e.g., POS category proportions, POS bigram entropy/diversity, and dependency-parse structure metrics. ⢠Conversational/discourse structure: e.g., turn-taking, question types, caregiver/child balance, overlap markers, repair/expansion cues. ⢠Mixture/composition metadata: e.g., number of component families, age-distribution statistics, and cross-family divergence features. ⢠Semantic and predictability proxies: e.g., adjacent-turn semantic similarity and normalized reference-LM perplexity features. ⢠Transcription/data-quality indicators: e.g., unintelligible markers, partial-word and non-linguistic token rates. We analyze how well the linguistic features predict model performance across the 25 BabyView experiments discussed in §3.2 for the four different model-scale combinations discussed in §3.3. We focus on three evaluation targets: Zorro, WordSim, and COMPS.222We exclude EWoK due to poor scaling behavior in our BabyView-scale domain. We run three complementary analyses for each model-target cell: (i) Spearman feature-target correlations, (i) LassoCV with standardized predictors, and (i) XGBoost regressors with feature-importance extraction. This provides a direct, method-agnostic view of which linguistic predictors recur across model-target settings. 3.6 Predicting Childrenās Vocabulary Outcomes We were also interested in evaluating connections between model performance and childrenās learning outcomes (see Portelance et al., 2023). One gold-standard measure for childrenās early vocabulary is the MacArthur-Bates Communicative Development Inventory (CDI; Fenson and others, 2007; Frank et al., 2021), a validated parent-report vocabulary checklist. The families in BabyView filled out CDI forms approximately once every 3 months over the course of the study. For each trained model, we computed the likelihood assigned to words appearing in the CDI vocabulary inventory. Specifically, we measured the average negative log-likelihood (NLL) of CDI words under each model by recursively splitting each example in our training corpora into 180-word contexts.333Required to fit them under the maximum sequence length of the corresponding models. This metric serves as a proxy for how well the training data supports learning developmentally relevant vocabulary: lower NLL (i.e., higher probability) indicates that the model assigns higher probability to CDI vocabulary items. We calculated the age of acquisition (AoA) for each CDI item using the expressive vocabulary data in the BabyView CDIs, fitting a Bayesian binomial regression to the data predicting word production from age. We used weakly informative priors (binterceptā¼ā(0,2.5)b_ intercept (0,2.5); bageā¼ā(0.3,0.1)b_ age (0.3,0.1)). We then calculated AoA as the negative of the intercept divided by the age slope; the AoA is the age at which 50% of children are expected to know a word. We then compared a base frequency regression and an NLL regression: the base regression predicted AoA as a function of log frequency, concreteness, and their interactions with lexical category, while the NLL regression substituted model-derived NLL for log frequency ā theoretically a more sensitive measure of processing difficulty than log frequency (Portelance et al., 2023). To understand possible biasāvariance tradeoffs from the training data, we considered three settings for obtaining NLL estimates: 1) averaging across the individual-family models, 2) the all-families models, and 3) the models trained on CHILDES (24.5M word English subset). We evaluated models in terms of their Akaike Information Criterion (AIC). 4 Results Figure 1: Broader performance scaling with data quantity across all four benchmarks. Each color denotes a separate dataset, with different markers showing different models. Gray curves show spline-smoothed average relationships across all data. Model Condition # of Fams Tokens TTR Zorro WordSim COMPS EWoK GPT-2 small Mean of fams 1 142,094 ā 55.19 ± 2.89 0.049 ± 0.007 50.09 ± 0.14 50.18 ± 0.15 GPT-2 small Top-1 1 725,600 0.026 58.09 ± 3.36 0.042 ± 0.025 50.19 ± 0.49 50.07 ± 0.36 GPT-2 small Top-2 2 727,986 0.025 59.58 ± 1.85 0.064 ± 0.022 50.37 ± 0.33 50.12 ± 0.13 GPT-2 small Top-5 5 727,883 0.026 56.70 ± 2.88 0.062 ± 0.008 50.22 ± 0.27 50.32 ± 0.05 GPT-2 small Top-10 10 735,309 0.027 57.14 ± 2.79 0.048 ± 0.020 50.22 ± 0.20 50.13 ± 0.21 GPT-2 small Top-20 20 734,234 0.026 56.59 ± 2.18 0.044 ± 0.015 50.03 ± 0.23 50.31 ± 0.29 GPT-2 small All fams 20 2,841,873 0.013 67.14 ± 0.78 0.154 ± 0.018 50.53 ± 0.34 50.10 ± 0.44 GPT-2 mini Mean of fams 1 142,094 ā 55.92 ± 2.85 0.031 ± 0.018 50.06 ± 0.15 50.04 ± 0.16 GPT-2 mini Top-1 1 725,600 0.026 57.54 ± 1.98 0.044 ± 0.018 50.10 ± 0.16 50.03 ± 0.26 GPT-2 mini Top-2 2 727,986 0.025 60.42 ± 2.09 0.084 ± 0.040 50.14 ± 0.17 50.08 ± 0.19 GPT-2 mini Top-5 5 727,883 0.026 58.34 ± 1.20 0.060 ± 0.027 50.20 ± 0.36 50.31 ± 0.64 GPT-2 mini Top-10 10 735,309 0.027 59.74 ± 1.48 0.056 ± 0.026 50.13 ± 0.08 49.86 ± 0.09 GPT-2 mini Top-20 20 734,234 0.026 58.66 ± 1.68 0.038 ± 0.011 49.91 ± 0.19 49.85 ± 0.59 GPT-2 mini All fams 20 2,841,873 0.013 62.71 ± 2.50 0.126 ± 0.025 50.41 ± 0.09 50.09 ± 0.25 GPT-BERT base Mean of fams 1 142,094 ā 57.77 ± 1.79 0.036 ± 0.024 50.00 ± 0.16 50.06 ± 0.11 GPT-BERT base Top-1 1 725,600 0.026 59.10 ± 1.58 0.068 ± 0.026 50.22 ± 0.27 50.31 ± 0.35 GPT-BERT base Top-2 2 727,986 0.025 60.17 ± 1.81 0.064 ± 0.011 49.93 ± 0.26 50.09 ± 0.17 GPT-BERT base Top-5 5 727,883 0.026 60.80 ± 2.80 0.082 ± 0.015 49.91 ± 0.21 49.76 ± 0.00 GPT-BERT base Top-10 10 735,309 0.027 59.35 ± 1.97 0.074 ± 0.005 50.21 ± 0.25 50.10 ± 0.15 GPT-BERT base Top-20 20 734,234 0.026 59.86 ± 2.45 0.076 ± 0.019 50.03 ± 0.42 50.12 ± 0.22 GPT-BERT base All fams 20 2,841,873 0.013 66.95 ± 2.09 0.166 ± 0.005 50.86 ± 0.24 49.95 ± 0.44 GPT-BERT small Mean of fams 1 142,094 ā 56.94 ± 1.38 0.049 ± 0.021 50.03 ± 0.13 49.98 ± 0.10 GPT-BERT small Top-1 1 725,600 0.026 54.37 ± 1.63 0.072 ± 0.016 50.10 ± 0.10 50.00 ± 0.11 GPT-BERT small Top-2 2 727,986 0.025 54.86 ± 1.17 0.056 ± 0.021 49.84 ± 0.11 50.05 ± 0.15 GPT-BERT small Top-5 5 727,883 0.026 57.08 ± 1.74 0.078 ± 0.018 49.98 ± 0.25 50.05 ± 0.10 GPT-BERT small Top-10 10 735,309 0.027 56.88 ± 1.14 0.058 ± 0.019 50.16 ± 0.34 50.01 ± 0.20 GPT-BERT small Top-20 20 734,234 0.026 56.76 ± 1.08 0.062 ± 0.015 49.96 ± 0.15 50.10 ± 0.36 GPT-BERT small All fams 20 2,841,873 0.013 63.58 ± 1.73 0.086 ± 0.015 50.54 ± 0.49 49.95 ± 0.18 Table 1: Major BabyView results (averaged across 5 seeds per experiment). Mean of fams reports mean and standard deviation across the 20 individual-family datasets. Pooled rows (Top-2/5/10/20) are proportional mixtures with seed-level variability. Top-1 is the largest single family; All fams is the all-families model trained on the concatenated corpus. Zorro, COMPS, and EWoK are percentages (higher is better); WordSim is Spearman correlation (higher is better). TTR is typeātoken ratio (unique word types divided by total tokens). Figure 2: Performance scaling of (A) BabyView experiments and (B) TinyDialogues experiments across all four benchmarks. Each color denotes a separate model; different markers show different experimental conditions. Lines of best fit show average linear relationships. 4.1 Scaling at Child-Scale Data Regimes Scaling within BabyView datasets. We first examine scaling behavior within BabyView itself. Each family dataset contains a different number of tokens, allowing us to study how model performance varies as a function of input size. Major results are shown in Table 1 and Figures 1 and 2A.444Explicit individual family results are in Tables 3 to 6 in Appendix D. Across benchmarks other than EWoK (world knowledge) which shows little to no signal, we observe a positive relationship between dataset size and model performance, but scaling slopes are strong for grammatical knowledge (Zorro) and weaker for WordSim and COMPS, which represent semantic knowledge. Larger datasets generally produce stronger models, although results were notably variable between families. Mixture vs single-family datasets. We next compared models trained on mixtures of families with those trained on individual families (Table 1). More families were not always monotonically better: at matched token budgets, some pooled mixtures outperform the largest individual family (Top-1), but the pattern was benchmark and architecture dependent. Importantly, adding families did not increase raw lexical diversity as measured by typeātoken ratio (TTR). This suggests that the benefits of pooling, when they occur, may arise less from adding many new word types and more from increased contextual, discourse, speaker, and/or interactional diversity. See §4.2 below for further analysis of these features. Broader scaling behavior. To extend our scaling curves, we trained our models on increasing amounts of the synthetic TinyDialogues (TD) data ranging from 10M to 200M tokens (Figure 2B). Across all four benchmarks, model performance improves ā¼ linearly with the logarithm of training tokens, consistent with classical scaling law behavior. We further compare to results on the 24.5M word English-subset of CHILDES, following Feng et al. (2024b), and a pretrained RoBERTa-base on 30B tokens (serving as a topline). As seen in Figure 1, all four benchmarks exhibit classical scaling trends, and scaling is notably different than with BabyView data. To quantify this trend, we fit linear models to each model and benchmark predicting performance as a function of dataset, log(tokens), and their interaction. Of the 16 models, all but four (GPT-BERT scaling on Zorro and WordSim) showed significant interactions, indicating different scaling curves for BabyView vs. TinyDialogues. One possibility is that the synthetic TinyDialogues data is richer and more variable in content than true child data, leading to stronger scaling. 4.2 Performance Variability Across Families Despite similar dataset sizes, models trained on different families showed noticeably different performance, as seen in Figure 2A. Although there was a positive relationship between token count and performance, families with comparable sizes produced models whose scores differed by several points, as did different mixtures of families. Relative to seed-level variability, between-experiment spread is comparable and can be larger for specific model-metric pairs. Thus, dataset size alone likely does not fully determine learning outcomes. Feature Top10 Meth. Spear. Lasso XGB Mean rank Mean |Ļ||Ļ| Mean |β||β| Mean imp. Bigram mutual information 9 3 3 1 5 5.56 0.72 0.153 0.060 Caregiver POS-tagged token count 7 3 4 1 2 4.29 0.74 0.016 0.062 Child POS-tagged token count 7 3 1 4 2 2.00 0.86 0.604 0.356 Child-to-caregiver semantic pair count 6 3 2 2 2 6.17 0.80 0.005 0.050 POS bigram entropy 6 3 1 1 4 3.33 0.56 0.263 0.069 Mean KL divergence from other datasets 5 3 3 1 1 3.60 0.71 0.500 0.053 Total parse-eligible utterances 5 3 1 1 3 3.40 0.85 0.012 0.149 Hapax ratio 5 3 1 1 3 6.60 0.48 0.264 0.067 Trigram entropy 5 2 4 0 1 5.00 0.74 0.000 0.075 Number of conversations 5 1 0 5 0 2.00 0.00 0.178 0.000 Table 2: Top linguistic predictors across statistical analyses, identified and sorted by Top10 hit. Top10 counts the number of times a feature appears in top-10 lists across all methods and experimental conditions. Meth. is the number of distinct methods (Spearman, Lasso, XGBoost) that select the feature at least once. Spear., Lasso, and XGB report method-specific top-10 counts. The following mean columns compute average across top-10 appearances. Mean rank is the average within-list rank when selected. Mean |Ļ||Ļ| is the average absolute Spearman correlation. Mean |β||β| is the average absolute lasso coefficient (linear effect size under regularization). Mean imp. is the average XGBoost feature importance (nonlinear contribution to predictive performance). Full results (all linguistic features) in Table LABEL:app_tab:ling_full_grouped. Table 2 summarizes the top linguistic features found to be most influential on downstream performance across the 25 BabyView experiments and 4 model-size combinations (see also Table 7 in Appendix F). The strongest predictors included distributional and syntactic features such as bigram mutual information, POS bigram entropy, and POS-tagged token coverage (caregiver and child). Further, divergence and coverage features (e.g., semantic pair coverage, mean KL from other datasets, parse-eligible utterance count, and number of conversations) were also related to downstream performance, suggesting that models exposed to broader and more consistently structured linguistic contexts learn more effectively. These variables indicate that multiple facets of input quality, composition, and coverage shape model performance, rather than dataset size alone. Figure 3: (A) Relationship between mean token NLL and AoA for CDI words by lexical category, for models trained on data from all families. (B) AIC of the regressions predicting AoA from mean token NLL. Dashed line indicates AIC of the base frequency model. 4.3 Connections to Child Vocabulary Outcomes We examined the relationship between model token likelihood and childrenās reported vocabulary outcomes on the CDI, using the age of acquisition of CDI words. While token NLL appeared to show some relationships with AoA (Figure 3A for models trained on data from all families), none of the regressions containing model NLL values attained a lower AIC than the base frequency model (Figure 3B). We note that there are moderate negative correlations between log frequency and mean token NLL (rā[ā.56,ā.42]rā[-.56,-.42]), suggesting that NLL values share some variance with log frequency but may be capturing additional noise. 5 Discussion We examined the scaling properties of SLMs trained on data from childrenās linguistic experiences, and ā inspired by childrenās language learning ā asked whether features of particular familiesā language environment lead to better scaling. Classical scaling laws appear to persist even in extremely small data regimes for grammaticality: model performance improved linearly in the logarithm of training tokens. Models scaled less well on semantic tasks, however, and there was almost no increase in performance on world knowledge for our BabyView experiments. Perhaps the restricted data from interactions in childrenās home environments is insufficient input for learning about the broader world. More diverse input data ā e.g., from books or experiences outside the home ā might play an important role for semantic generalization and world knowledge acquisition. Variability across individual language environments was substantial. Models trained on different familiesā input achieved markedly different performance levels even when dataset sizes were similar. This finding parallels observations in language development research showing that children experience highly variable language environments, and that these environments relate to their learning outcomes (Hart et al., 1997; Weisleder and Fernald, 2013; Coffey and Snedeker, 2025). Our linguistic feature analysis suggests that multiple factors, including distributional structure, syntactic composition and coverage, and semantic coverage, play an important role in shaping model performance across multiple tasks, aligning with prior work showing that modifying the structure and composition of input representations can significantly affect LM generation quality (e.g., Feng et al., 2021). This again mirrors findings from child language development that emphasize quality over pure quantity of input (Anderson et al., 2021; Masek et al., 2021; Rowe, 2012; Rowe and Snow, 2020). This also aligns with our above finding, suggesting that more diverse and high-quality input data is likely necessary for effective LM learning, especially at such small scales. Finally, the relationship between model CDI likelihood and childrenās vocabulary outcomes suggests that some properties of child-directed input may influence both machine and human language learning. While these results are preliminary, they point toward the possibility of using LMs as computational probes for studying the structure of individual childrenās linguistic environments, much as has been done with larger scale data already (e.g., Portelance et al., 2023). 6 Conclusion Together, these findings highlight the importance of considering not only the quantity of language input but also its composition, structure, and coverage. We have only scratched the surface of this goal here: future work should move beyond transcribed speech to consider the grounded, multimodal nature of language exposure. Further, from a psychological perspective, the generalizability of our work here is limited by the relatively small sample size we examined; while the BabyView corpus is the largest of its kind, it still represents only one language (English) and a small, selected subset of families. Nevertheless, we hope that our work sets the stage for future investigations of childrenās language learning environments using language models. Understanding what makes language data efficient for learning may ultimately inform both the design of more data-efficient language models and scientific theories of human language acquisition. Acknowledgments We gratefully acknowledge Modal for providing a portion of the compute resources that enabled this work. References N. J. Anderson, S. A. Graham, H. Prime, J. M. Jenkins, and S. Madigan (2021) Linking quality and quantity of parental linguistic input to child language skills: a meta-analysis. Child Development 92 (2), p. 484ā501. Cited by: §1, §2, §5. A. Brill (2024) Neural scaling laws rooted in the data distribution. External Links: 2412.07942, Link Cited by: §2. E. Bruni, G. Boleda, M. Baroni, and N. Tran (2012) Distributional semantics in technicolor. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), H. Li, C. Lin, M. Osborne, G. G. Lee, and J. C. Park (Eds.), Jeju Island, Korea, p. 136ā145. External Links: Link Cited by: Appendix C. L. G. G. Charpentier and D. Samuel (2024) GPT or bert: why not both?. arXiv preprint arXiv:2410.24159. Cited by: §1, §3.3. J. R. Coffey and J. Snedeker (2025) How strong is the relationship between caregiver speech and language development? a meta-analysis. Journal of Child Language, p. 1ā36. Cited by: §2, §5. I. Constantinescu, T. Pimentel, R. Cotterell, and A. Warstadt (2025) Investigating critical period effects in language acquisition through neural language models. Transactions of the Association for Computational Linguistics 13, p. 96ā120. External Links: Link, Document Cited by: §2. R. Eldan and Y. Li (2023) Tinystories: how small can language models be and still speak coherent english?. arXiv preprint arXiv:2305.07759. Cited by: §3.1. S. Feng, S. Prabhumoye, K. Kong, D. Su, M. Patwary, M. Shoeybi, and B. Catanzaro (2024a) Maximize your dataās potential: enhancing llm accuracy with two-phase pretraining. External Links: 2412.15285, Link Cited by: §2. S. Y. Feng, N. D. Goodman, and M. C. Frank (2024b) Is child-directed speech effective training data for language models?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 22055ā22071. External Links: Link, Document Cited by: Appendix A, Appendix A, §2, §3.1, §3.4, §4.1. S. Y. Feng, J. Huynh, C. P. Narisetty, E. Hovy, and V. Gangal (2021) SAPPHIRE: approaches for enhanced concept-to-text generation. In Proceedings of the 14th International Conference on Natural Language Generation, A. Belz, A. Fan, E. Reiter, and Y. Sripada (Eds.), Aberdeen, Scotland, UK, p. 212ā225. External Links: Link, Document Cited by: §5. L. Fenson et al. (2007) MacArthur-bates communicative development inventories. Cited by: §3.6. L. Finkelstein, E. Gabrilovich, Y. Matias, E. Rivlin, Z. Solan, G. Wolfman, and E. Ruppin (2001) Placing search in context: the concept revisited. ACM Transactions on Information Systems - TOIS 20, p. 406ā414. External Links: Document Cited by: Appendix C. M. C. Frank, M. Braginsky, D. Yurovsky, and V. A. Marchman (2021) Variability and consistency in early language learning: the wordbank project. MIT Press. Cited by: §2, §3.6. M. C. Frank (2023a) Bridging the data gap between children and large language models. Trends in Cognitive Sciences. Cited by: §1. M. C. Frank (2023b) Openly accessible llms can help us to understand human cognition. Nature Human Behaviour 7 (11), p. 1825ā1827. Cited by: §2. L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy (2020) The pile: an 800gb dataset of diverse text for language modeling. External Links: 2101.00027, Link Cited by: §2. D. Gerz, I. VuliÄ, F. Hill, R. Reichart, and A. Korhonen (2016) SimVerb-3500: a large-scale evaluation set of verb similarity. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, J. Su, K. Duh, and X. Carreras (Eds.), Austin, Texas, p. 2173ā2182. External Links: Link, Document Cited by: Appendix C. B. Hart, T. R. Risley, and J. R. Kirby (1997) Meaningful differences in the everyday experience of young american children. Canadian Journal of Education 22 (3), p. 323. Cited by: §2, §5. T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, C. Hallacy, B. Mann, A. Radford, A. Ramesh, N. Ryder, D. M. Ziegler, J. Schulman, D. Amodei, and S. McCandlish (2020) Scaling laws for autoregressive generative modeling. External Links: 2010.14701, Link Cited by: §2. F. Hill, R. Reichart, and A. Korhonen (2015) SimLex-999: evaluating semantic models with (genuine) similarity estimation. Computational Linguistics 41 (4), p. 665ā695. External Links: Link, Document Cited by: Appendix C. J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, S. Bowman, et al. (2022) Training compute-optimal large language models. In Advances in Neural Information Processing Systems, Vol. 35, p. 300ā313. External Links: Link Cited by: §2. J. Hu, A. W. M. Tan, S. Y. Feng, and M. C. Frank (2025) Language production is harder than comprehension for children and language models. In Proceedings of the Annual Meeting of the Cognitive Science Society, Vol. 47. External Links: Link Cited by: §2. M. Y. Hu, A. Mueller, C. Ross, A. Williams, T. Linzen, C. Zhuang, R. Cotterell, L. Choshen, A. Warstadt, and E. G. Wilcox (2024) Findings of the second babylm challenge: sample-efficient pretraining on developmentally plausible corpora. In The 2nd BabyLM Challenge at the 28th Conference on Computational Natural Language Learning, p. 1ā21. Cited by: §1, §2, §3.4. P. A. Huebner, E. Sulem, F. Cynthia, and D. Roth (2021) BabyBERTa: learning more grammar with small-scale child-directed language. In Proceedings of the 25th conference on computational natural language learning, p. 624ā646. Cited by: Appendix A, §2, §3.1, §3.4. A. A. Ivanova, A. Sathe, B. Lipkin, U. U. Kumar, S. Radkani, T. H. Clark, C. Kauf, J. Hu, R. T. Pramod, G. Grand, V. C. Paulun, M. Ryskina, E. Akyürek, E. G. Wilcox, N. Rashid, L. Choshen, R. Levy, E. Fedorenko, J. Tenenbaum, and J. Andreas (2025) Elements of world knowledge (ewok): a cognition-inspired framework for evaluating basic world knowledge in language models. Transactions of the Association for Computational Linguistics 13, p. 1245ā1270. External Links: ISSN 2307-387X, Document, Link, https://direct.mit.edu/tacl/article-pdf/doi/10.1162/TACL.a.38/2557964/tacl.a.38.pdf Cited by: §3.4. J. Kallini, I. Papadimitriou, R. Futrell, K. Mahowald, and C. Potts (2024) Mission: impossible language models. arXiv preprint arXiv:2401.06416. Cited by: §2, §3.3. J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei (2020) Scaling laws for neural language models. External Links: 2001.08361, Link Cited by: §2. E. Kidd and S. Donnelly (2020) Individual differences in first language acquisition. Annual review of linguistics 6 (1), p. 319ā340. Cited by: §2. T. Kunze, M. MĆ©tais, H. Titeux, L. Elbert, J. Coffey, E. Dupoux, A. Cristia, and M. Lavechin (2025) Challenges in automated processing of speech from child wearables: the case of voice type classifier. arXiv preprint arXiv:2506.11074. Cited by: §3.1. B. Long, R. Z. Sparks, V. Xiang, S. Stojanov, Z. Yin, G. E. Keene, A. W. Tan, S. Y. Feng, C. Zhuang, V. A. Marchman, et al. (2024) The babyview dataset: high-resolution egocentric videos of infantsā and young childrenās everyday experiences. arXiv preprint arXiv:2406.10447. Cited by: §1, §3.1. S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y. Tay, D. Zhou, Q. V. Le, B. Zoph, J. Wei, and A. Roberts (2023) The flan collection: designing data and methods for effective instruction tuning. In Proceedings of the 40th International Conference on Machine Learning, ICMLā23. Cited by: §2. B. MacWhinney (2014) The childes project: tools for analyzing talk, volume i: transcription format and programs. Psychology Press. Cited by: §1, §3.1. R. D. Martinez, H. McGovern, Z. Goriely, C. Davis, A. Caines, P. Buttery, and L. Beinborn (2023) CLIMB ā curriculum learning for infant-inspired model building. In Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning, A. Warstadt, A. Mueller, L. Choshen, E. Wilcox, C. Zhuang, J. Ciro, R. Mosquera, B. Paranjabe, A. Williams, T. Linzen, and R. Cotterell (Eds.), Singapore, p. 112ā127. External Links: Link, Document Cited by: §1. L. R. Masek, A. G. Ramirez, B. T. McMillan, K. Hirsh-Pasek, and R. M. Golinkoff (2021) Beyond counting words: a paradigm shift for the study of language acquisition. Child Development Perspectives 15 (4), p. 274ā280. Cited by: §1, §2, §5. K. Misra and K. Mahowald (2024) Language models learn rare phenomena from less rare phenomena: the case of the missing aanns. External Links: 2403.19827 Cited by: §2, §3.3. K. Misra, J. Rayz, and A. Ettinger (2023) COMPS: conceptual minimal pair sentences for testing robust property knowledge and its inheritance in pre-trained language models. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, A. Vlachos and I. Augenstein (Eds.), Dubrovnik, Croatia, p. 2928ā2949. External Links: Link, Document Cited by: §3.4. N. Muennighoff, A. M. Rush, B. Barak, T. L. Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. Raffel (2023) Scaling data-constrained language models. In Thirty-seventh Conference on Neural Information Processing Systems, External Links: Link Cited by: §2. R. Pandey (2024) Gzip predicts data-dependent scaling laws. External Links: 2405.16684, Link Cited by: §2. E. Portelance, Y. Duan, M. C. Frank, and G. Lupyan (2023) Predicting age of acquisition for childrenās early vocabulary in five languages using language model surprisal. Cognitive Science 47 (9), p. e13334. Cited by: §3.6, §3.6, §5. Y. Qin, W. Wang, and B. M. Lake (2024) A systematic investigation of learnability from single child linguistic input. In Proceedings of the 46th Annual Conference of the Cognitive Science Society, Cited by: §1, §2, §3.3. A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever (2023) Robust speech recognition via large-scale weak supervision. In International Conference on Machine Learning, p. 28492ā28518. Cited by: §3.1. A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al. (2019) Language models are unsupervised multitask learners. OpenAI blog 1 (8), p. 9. Cited by: §3.3. M. L. Rowe and C. E. Snow (2020) Analyzing input quality along three dimensions: interactive, linguistic, and conceptual.. Journal of child language 47 (1). Cited by: §1, §2, §5. M. L. Rowe (2012) A longitudinal investigation of the role of quantity and quality of child-directed speech in vocabulary development. Child development 83 (5), p. 1762ā1774. Cited by: §1, §2, §5. H. Rubenstein and J. B. Goodenough (1965) Contextual correlates of synonymy. In Communications of the ACM (CACM) 8 (10), Cited by: Appendix C. B. Sorscher, R. Geirhos, S. Shekhar, S. Ganguli, and A. Morcos (2022) Beyond neural scaling laws: beating power law scaling via data pruning. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35, p. 19523ā19536. External Links: Link Cited by: §2. K. Tirumala, D. Simig, A. Aghajanyan, and A. Morcos (2023) D4: improving llm pretraining via document de-duplication and diversification. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, p. 53983ā53995. External Links: Link Cited by: §2. W. Wang, W. K. Vong, N. Kim, and B. M. Lake (2023) Finding structure in one childās linguistic experience. Cognitive science 47 (6), p. e13305. Cited by: §1. A. Warstadt, A. Mueller, L. Choshen, E. Wilcox, C. Zhuang, J. Ciro, R. Mosquera, B. Paranjabe, A. Williams, T. Linzen, et al. (2023) Findings of the babylm challenge: sample-efficient pretraining on developmentally plausible corpora. In Proceedings of the BabyLM Challenge at the 27th Conference on Computational Natural Language Learning, Cited by: §1, §1, §2. A. Weisleder and A. Fernald (2013) Talking to children matters: early language experience strengthens processing and builds vocabulary. Psychological science 24 (11), p. 2143ā2152. Cited by: §2, §5. Y. Zhang, A. Warstadt, X. Li, and S. R. Bowman (2021) When do you need billions of words of pretraining data?. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C. Zong, F. Xia, W. Li, and R. Navigli (Eds.), Online, p. 1112ā1125. External Links: Link, Document Cited by: §2. C. Zhuang, E. Fedorenko, and J. Andreas (2023) Visual grounding helps learn word meanings in low-data regimes. arXiv preprint arXiv:2310.13257. Cited by: Appendix C, §3.4. Appendix A Dataset Examples & Preprocessing Details Following Feng et al. (2024b), we preprocess training examples in both BabyView and TinyDialogues in the same way. Double asterisks surround speaker labels, double newline tokens separate utterances, and an end-of-text token marks the end of the conversation. This format was consistent across all conversations in both BabyView and TinyDialogues. BabyView example: **MOT**: Put it in the oven for baby and me. **MOT**: Oh, hold on. **MOT**: Letās try to put it on. **CHI**: Whoopsie. **MOT**: Whoopsie. **OCHI**: Hey, Rosa, look at me. **MOT**: Look at me. (ā¦) <|endoftext|> TinyDialogues example: **Teacher**: Alright, everyone, itās time to clean up! **Child**, can you please help me by putting the crayons back in the box? **Child**: Yes! I can do that. The box is empty so Iāl fill it up! **Teacher**: Thank you, thatās very helpful. Make sure the lids are on tight so they donāt dry out. **Child**: I did it! Look, theyāre all inside now. (ā¦) <|endoftext|> Training data was set up so that each line corresponded to a single conversation. Removing the childās speech from the dialogue would create incoherent training data that lacked context; thus, consistent with previous work leveraging transcripts of childrenās language environment, we include both speaker labels and the childās own speech in our training data (Huebner et al., 2021; Feng et al., 2024b) . Appendix B Further Training & Compute Details We trained a separate tokenizer for each model on BāVEānāgBV_Eng, TinyDialogues, and CHILDES. We use a vocab size of 50k for GPT-2, 16k for GPT-BERT base, and 8k for GPT-BERT small. For each training run, the best checkpoint was selected based on lowest validation loss. We pretrain GPT-2 from scratch, using a linear LR scheduler with no warmup, varying batch sizes per GPU (4 to 64 depending on the GPU and experiment), and Adam optimizer with β=(0.9,0.999)β=(0.9,0.999) and ϵ=1āeā08ε=1e-08. During training, GPT-2 processes data in 1024-token chunks. To allow the models to distinguish between individual conversations and training examples, we add an end-of-text token to the end of each conversation. We train GPT-2 small for up to 20 epochs for all experiments, and GPT-2 mini for up to 50 epochs for BabyView experiments and 20 epochs for TinyDialogues and CHILDES. Our GPT-2 experiments are based on HuggingFaceās causal LM training code, with modifications for our experiments. We pretrain GPT-BERT from scratch, using a cosine LR scheduler with 0.1 weight decay. On the BabyView experiments, we use a sequence length which increases throughout training (128 to 256 after 70% of training and 512 after 90% of training) while the batch size decreases accordingly with each sequence length increase (down to 1/2 and 1/4 ā from 2048 for small and 1024 for base down to 512 and 256, respectively). For our TinyDialogues and CHILDES training, we keep sequence length (128) and batch size (512 for small, 256 for base) fixed throughout training. Further, due to GPT-BERTās hybrid objective that mixes autoregressive next-token prediction and masked-token prediction, we try different causal:mask ratios (e.g., 1:1 means half causal and half masked training workers), while the implementation maps this to a masked-worker fraction under GPU divisibility constraints. We search through 0:1, 1:7, 1:3, 1:1, 3:1, 7:1, 1:0 ratios, and settle on 1:1 for our final training experiments due to the most stable evaluation and scaling behavior. Our GPT-BERT experiments are based on the original authorsā GitHub repo and code555https://github.com/ltgoslo/gpt-bert, with modifications for our experiments. We train 5 seeds 42,0,123,321,9000\42,0,123,321,9000\ of each model for every BāVEānāgBV_Eng experiment, and a single seed (either 42 or 0) for TinyDialogues and CHILDES. We searched through different values of the learning rate (LR) for GPT-2 training. Specifically, LāR=1āeā06,5āeā06,1āeā05,5āeā05,1āeā04,3āeā04,5āeā04,1āeā03LR=\1e-06,5e-06,1e-05,5e-05,1e-04,3e-04,5e-04,1e-03\. Through initial experiments, we found that LāR=1āeā04LR=1e-04 led to the best convergence behavior for small and LāR=3āeā04LR=3e-04 for mini, and used that for all our training experiments. We do the same for GPT-BERT, searching through LāR=3āeā5,5āeā5,1āeā4,5āeā4,1āeā3,3āeā3,5āeā3,1āeā2,3āeā2,5āeā2LR=\3e-5,5e-5,1e-4,5e-4,1e-3,3e-3,5e-3,1e-2,3e-2,5e-2\, and choose LāR=1āeā02LR=1e-02 for both model sizes for BabyView experiments, and LāR=5āeā3LR=5e-3 for base and LāR=1āeā2LR=1e-2 for small for TinyDialogues and CHILDES. Due to GPT-BERTās architecture, we report three ways (inference modes) of conducting Zorro evaluation of GPT-BERT: causal, bidirectional, and fused. We find that results can vary depending on the inference mode chosen and the causal:mask ratio used for training. Causal applies an autoregressive triangular attention mask, bidirectional uses full attention, and fused computes both passes and sums their logits before scoring. For each set of experiments, we report the Zorro inference mode that results in the most stable evaluation and scaling behavior. Our experiments were run on varying GPUs. This included a single RTX 3090TI and up to eight A40s, A100s, H100s, and H200s. Training time varied by the type and number of GPUs used and the particular experiment. Appendix C Word Similarity Benchmark Details Following Zhuang et al. (2023), we extract word embeddings from the hidden layers of each model and compute pairwise cosine similarities. The best layer of each model is chosen. We average results across several word similarity benchmarks including RG-65 (Rubenstein and Goodenough, 1965), WordSim-353 (Finkelstein et al., 2001), SimLex-999 (Hill et al., 2015), SimVerb-3500 (Gerz et al., 2016), and MEN (MTest-3000) (Bruni et al., 2012). Family Age (months) Tokens Zorro WordSim COMPS EWoK BV-fam-S00400001 31 725,600 58.09 ± 3.36 0.042 ± 0.025 50.19 ± 0.49 50.07 ± 0.36 BV-fam-S00510002 11 569,788 55.95 ± 2.74 0.054 ± 0.015 50.05 ± 0.42 50.14 ± 0.20 BV-fam-S00360001 8 257,817 59.53 ± 1.83 0.048 ± 0.022 50.07 ± 0.15 50.17 ± 0.30 BV-fam-S00220001 5 242,954 58.02 ± 2.19 0.058 ± 0.029 50.12 ± 0.34 50.18 ± 0.36 BV-fam-S00400002 28 181,317 57.95 ± 2.19 0.054 ± 0.031 50.07 ± 0.35 50.42 ± 0.80 BV-fam-S00320002 7 171,656 55.98 ± 2.31 0.028 ± 0.024 50.25 ± 0.32 50.21 ± 0.27 BV-fam-S00370001 8 167,136 58.28 ± 1.92 0.044 ± 0.023 50.02 ± 0.27 49.95 ± 0.20 BV-fam-S00430001 31 144,253 58.00 ± 2.25 0.056 ± 0.030 50.01 ± 0.24 50.39 ± 0.89 BV-fam-S00510001 11 131,120 55.61 ± 3.61 0.050 ± 0.023 50.19 ± 0.19 50.13 ± 0.31 BV-fam-S00720001 38 100,807 58.74 ± 1.71 0.046 ± 0.049 50.41 ± 0.41 50.03 ± 0.35 BV-fam-S00490001 27 40,835 54.38 ± 5.82 0.044 ± 0.032 49.88 ± 0.23 50.06 ± 0.24 BV-fam-S00230001 22 26,501 54.78 ± 4.63 0.054 ± 0.030 50.22 ± 0.16 50.24 ± 0.30 BV-fam-S00440001 26 21,313 52.86 ± 6.37 0.054 ± 0.017 50.20 ± 0.20 50.24 ± 0.17 BV-fam-S00360002 8 15,006 53.35 ± 2.22 0.054 ± 0.021 49.90 ± 0.51 50.33 ± 0.57 BV-fam-S00460001 11 14,618 53.52 ± 4.27 0.056 ± 0.021 49.96 ± 0.41 50.28 ± 0.12 BV-fam-S00350002 8 12,614 51.48 ± 2.89 0.052 ± 0.024 50.06 ± 0.53 50.41 ± 0.05 BV-fam-S00550001 31 10,952 53.73 ± 3.18 0.054 ± 0.023 49.96 ± 0.53 50.07 ± 0.42 BV-fam-S00340002 7 3,457 52.90 ± 2.86 0.048 ± 0.028 50.06 ± 0.20 50.28 ± 0.29 BV-fam-S00350001 25 2,959 51.45 ± 4.72 0.050 ± 0.014 49.84 ± 0.32 49.85 ± 0.35 BV-fam-S01010001 23 1,170 49.13 ± 5.16 0.042 ± 0.016 50.23 ± 0.48 50.17 ± 0.50 Table 3: Individual-family BabyView results for GPT-2 small with seed-level variability, sorted by training tokens. Age is months at onboarding time of that familyās given child. Family Age (months) Tokens Zorro WordSim COMPS EWoK BV-fam-S00400001 31 725,600 57.54 ± 1.98 0.044 ± 0.018 50.10 ± 0.16 50.03 ± 0.26 BV-fam-S00510002 11 569,788 57.65 ± 2.52 0.054 ± 0.027 49.97 ± 0.20 49.91 ± 0.10 BV-fam-S00360001 8 257,817 59.29 ± 0.95 0.032 ± 0.033 50.13 ± 0.29 49.94 ± 0.36 BV-fam-S00220001 5 242,954 59.43 ± 1.08 0.058 ± 0.018 49.74 ± 0.45 50.19 ± 0.30 BV-fam-S00400002 28 181,317 56.25 ± 3.03 0.014 ± 0.024 50.07 ± 0.32 50.16 ± 0.25 BV-fam-S00320002 7 171,656 56.49 ± 2.74 0.014 ± 0.019 50.42 ± 0.27 50.17 ± 0.38 BV-fam-S00370001 8 167,136 57.59 ± 1.09 0.030 ± 0.019 50.15 ± 0.26 49.89 ± 0.12 BV-fam-S00430001 31 144,253 58.65 ± 2.60 0.038 ± 0.022 50.01 ± 0.24 50.17 ± 0.31 BV-fam-S00510001 11 131,120 56.36 ± 1.33 0.052 ± 0.025 50.09 ± 0.25 50.09 ± 0.36 BV-fam-S00720001 38 100,807 59.66 ± 1.68 0.010 ± 0.022 50.02 ± 0.08 49.93 ± 0.66 BV-fam-S00490001 27 40,835 55.16 ± 3.38 0.020 ± 0.035 50.04 ± 0.38 50.05 ± 0.21 BV-fam-S00230001 22 26,501 57.34 ± 2.53 0.020 ± 0.027 49.90 ± 0.49 49.99 ± 0.15 BV-fam-S00440001 26 21,313 56.53 ± 3.06 -0.004 ± 0.024 50.03 ± 0.29 50.24 ± 0.27 BV-fam-S00360002 8 15,006 53.29 ± 2.73 0.022 ± 0.022 50.19 ± 0.44 50.06 ± 0.42 BV-fam-S00460001 11 14,618 56.02 ± 4.22 0.000 ± 0.029 49.95 ± 0.34 49.75 ± 0.43 BV-fam-S00350002 8 12,614 53.52 ± 2.64 0.034 ± 0.041 50.30 ± 0.35 50.07 ± 0.13 BV-fam-S00550001 31 10,952 55.75 ± 2.61 0.050 ± 0.025 49.87 ± 0.33 50.20 ± 0.25 BV-fam-S00340002 7 3,457 50.90 ± 3.88 0.042 ± 0.029 50.06 ± 0.36 50.27 ± 0.25 BV-fam-S00350001 25 2,959 50.81 ± 1.67 0.050 ± 0.017 50.17 ± 0.32 49.97 ± 0.58 BV-fam-S01010001 23 1,170 50.21 ± 4.74 0.040 ± 0.024 49.94 ± 0.30 49.69 ± 0.50 Table 4: Individual-family BabyView results for GPT-2 mini with seed-level variability, sorted by training tokens. Age is months at onboarding time of that familyās given child. Family Age (months) Tokens Zorro WordSim COMPS EWoK BV-fam-S00400001 31 725,600 59.10 ± 1.58 0.068 ± 0.026 50.22 ± 0.27 50.31 ± 0.35 BV-fam-S00510002 11 569,788 57.13 ± 1.69 0.062 ± 0.013 50.17 ± 0.20 50.13 ± 0.17 BV-fam-S00360001 8 257,817 58.26 ± 0.85 0.058 ± 0.015 50.18 ± 0.29 50.08 ± 0.27 BV-fam-S00220001 5 242,954 59.88 ± 1.82 0.056 ± 0.009 49.93 ± 0.23 50.10 ± 0.17 BV-fam-S00400002 28 181,317 57.57 ± 1.58 0.054 ± 0.015 49.82 ± 0.28 49.80 ± 0.27 BV-fam-S00320002 7 171,656 56.68 ± 0.94 0.026 ± 0.011 50.24 ± 0.34 50.14 ± 0.35 BV-fam-S00370001 8 167,136 58.51 ± 1.41 0.062 ± 0.022 50.24 ± 0.26 50.00 ± 0.30 BV-fam-S00430001 31 144,253 56.77 ± 1.13 0.046 ± 0.021 50.00 ± 0.26 50.07 ± 0.09 BV-fam-S00510001 11 131,120 57.68 ± 2.73 0.052 ± 0.031 50.07 ± 0.22 50.03 ± 0.45 BV-fam-S00720001 38 100,807 58.40 ± 1.95 0.020 ± 0.016 49.98 ± 0.17 49.97 ± 0.18 BV-fam-S00490001 27 40,835 58.56 ± 1.69 0.030 ± 0.019 49.84 ± 0.04 49.90 ± 0.25 BV-fam-S00230001 22 26,501 59.63 ± 1.46 0.028 ± 0.018 49.77 ± 0.20 50.00 ± 0.32 BV-fam-S00440001 26 21,313 57.72 ± 1.90 -0.018 ± 0.008 49.80 ± 0.38 50.08 ± 0.34 BV-fam-S00360002 8 15,006 53.75 ± 1.85 0.004 ± 0.011 50.10 ± 0.27 49.97 ± 0.30 BV-fam-S00460001 11 14,618 59.10 ± 1.34 0.012 ± 0.008 50.07 ± 0.32 50.18 ± 0.28 BV-fam-S00350002 8 12,614 55.67 ± 1.19 0.042 ± 0.016 49.81 ± 0.41 50.12 ± 0.15 BV-fam-S00550001 31 10,952 57.29 ± 2.22 0.050 ± 0.017 50.04 ± 0.28 50.10 ± 0.41 BV-fam-S00340002 7 3,457 61.83 ± 1.23 0.052 ± 0.015 49.85 ± 0.20 50.06 ± 0.26 BV-fam-S00350001 25 2,959 56.70 ± 2.17 0.006 ± 0.009 49.86 ± 0.33 49.97 ± 0.17 BV-fam-S01010001 23 1,170 55.19 ± 2.94 0.014 ± 0.018 49.97 ± 0.33 50.21 ± 0.34 Table 5: Individual-family BabyView results for GPT-BERT base with seed-level variability, sorted by training tokens. Age is months at onboarding time of that familyās given child. Family Age (months) Tokens Zorro WordSim COMPS EWoK BV-fam-S00400001 31 725,600 54.37 ± 1.63 0.072 ± 0.016 50.10 ± 0.10 50.00 ± 0.11 BV-fam-S00510002 11 569,788 55.65 ± 0.85 0.068 ± 0.047 50.12 ± 0.33 49.96 ± 0.20 BV-fam-S00360001 8 257,817 57.97 ± 1.34 0.066 ± 0.021 50.22 ± 0.23 49.88 ± 0.39 BV-fam-S00220001 5 242,954 57.24 ± 1.12 0.074 ± 0.018 49.98 ± 0.36 50.17 ± 0.34 BV-fam-S00400002 28 181,317 59.01 ± 1.77 0.058 ± 0.016 49.77 ± 0.23 50.06 ± 0.28 BV-fam-S00320002 7 171,656 56.21 ± 1.77 0.056 ± 0.019 50.07 ± 0.27 49.99 ± 0.26 BV-fam-S00370001 8 167,136 57.87 ± 0.77 0.066 ± 0.023 50.13 ± 0.25 50.01 ± 0.10 BV-fam-S00430001 31 144,253 56.96 ± 1.25 0.064 ± 0.018 50.08 ± 0.24 50.05 ± 0.14 BV-fam-S00510001 11 131,120 58.44 ± 2.12 0.072 ± 0.023 50.03 ± 0.16 50.05 ± 0.21 BV-fam-S00720001 38 100,807 57.60 ± 1.36 0.050 ± 0.029 49.95 ± 0.21 50.13 ± 0.29 BV-fam-S00490001 27 40,835 56.85 ± 0.92 0.032 ± 0.029 50.01 ± 0.26 50.06 ± 0.23 BV-fam-S00230001 22 26,501 57.69 ± 2.79 0.038 ± 0.025 50.05 ± 0.24 49.99 ± 0.13 BV-fam-S00440001 26 21,313 55.95 ± 2.04 -0.004 ± 0.015 49.77 ± 0.26 49.94 ± 0.20 BV-fam-S00360002 8 15,006 54.91 ± 1.44 0.036 ± 0.017 50.14 ± 0.20 49.93 ± 0.26 BV-fam-S00460001 11 14,618 56.60 ± 2.05 0.062 ± 0.023 50.11 ± 0.19 49.92 ± 0.15 BV-fam-S00350002 8 12,614 56.14 ± 1.43 0.026 ± 0.030 50.13 ± 0.34 49.80 ± 0.17 BV-fam-S00550001 31 10,952 58.51 ± 2.10 0.034 ± 0.013 49.88 ± 0.28 49.87 ± 0.41 BV-fam-S00340002 7 3,457 59.41 ± 1.92 0.052 ± 0.022 49.87 ± 0.19 50.01 ± 0.19 BV-fam-S00350001 25 2,959 56.17 ± 3.85 0.028 ± 0.011 49.99 ± 0.33 50.04 ± 0.17 BV-fam-S01010001 23 1,170 55.31 ± 0.93 0.032 ± 0.028 50.18 ± 0.25 49.78 ± 0.06 Table 6: Individual-family BabyView results for GPT-BERT small with seed-level variability, sorted by training tokens. Age is months at onboarding time of that familyās given child. Appendix D Individual BabyView Family Experiment Results See Tables 3 to 6 for BabyView evaluation results of the individual family models. Model Target α NZ feat. Lasso R2R^2 Lasso RMSE XGB R2R^2 XGB RMSE GPT-2 Mini Zorro 0.0488 21 0.657 1.784 0.395 2.369 GPT-2 Mini WordSim 0.0026 13 -0.208 0.029 -0.166 0.029 GPT-2 Mini COMPS 0.0668 0 -0.144 0.166 -1.076 0.223 GPT-2 Small Zorro 0.5230 6 0.289 2.974 0.172 3.208 GPT-2 Small WordSim 0.0173 1 -1.047 0.031 -1.943 0.037 GPT-2 Small COMPS 0.0713 3 -1.479 0.257 -0.601 0.207 GPT-BERT Base Zorro 0.3479 8 -0.164 2.675 -0.415 2.949 GPT-BERT Base WordSim 0.0022 10 0.456 0.026 0.137 0.032 GPT-BERT Base COMPS 0.0191 13 -0.117 0.236 -0.683 0.290 GPT-BERT Small Zorro 1.0621 0 -0.954 2.558 -0.386 2.154 GPT-BERT Small WordSim 0.0028 10 0.109 0.019 0.182 0.018 GPT-BERT Small COMPS 0.0707 1 -0.159 0.171 -0.592 0.200 Table 7: Cross-validated predictive performance of lasso regression and XGBoost models for each modelātarget pair. α denotes the lasso regularization strength selected via cross-validation. NZ feat. is the number of features with nonzero coefficients in the fitted lasso model, reflecting model sparsity. Lasso R2R^2 and XGB R2R^2 report the proportion of variance explained on held-out data (higher is better), while Lasso RMSE and XGB RMSE report root mean squared error (lower is better). Together, these metrics characterize both predictive performance and model complexity across methods. Appendix E LLM Usage Disclosure We used LLMs (e.g., GPT-5) to assist with parts of the coding process and limited aspects of paper preparation, including LaTeX table formatting and minor editing for clarity and grammar. All outputs were carefully reviewed to ensure accuracy and appropriateness. Appendix F Linguistic Analysis: All Linguistic Features & Results See Table 7 for cross-validated Lasso and XGBoost predictive performance of our models on each target benchmark (Zorro, WordSim, COMPS). We see that our linguistic features can explain Zorro performance reasonably well for some models, but results are inconsistent across architectures. WordSim is only partially explained (primarily for GPT-BERT models), while COMPS shows low between-experiment variance and remains difficult to explain from the current feature set. See Table LABEL:app_tab:ling_feature_catalog_detailed for a full catalog (list and description) of all 175 linguistic features we considered for our analysis, and Table LABEL:app_tab:ling_full_grouped for statistical analysis results of all 175 linguistic features. Table 8: Comprehensive catalog of the 175 linguistic features considered in our BabyView analysis. The table is grouped by feature family and provides a one-sentence description of each feature. Note that all 175 features were kept after restricting to numeric variables and applying missingness (ā¤0.4⤠0.4 across the 25 experiments) and zero-variance filtering. Hence, all 175 features were used in the final statistical analyses (Spearman, lasso, and XGBoost). Feature name Brief description Lexical / Distributional / Scale avg_conversation_len_tokens This feature measures the average number of tokens per conversation, capturing how much language is typically contained in a single interaction. avg_sentence_len_tokens This feature measures the average number of tokens per sentence, providing a simple proxy for sentence-level length and complexity. avg_utterance_len_tokens This feature measures the average number of tokens per utterance, summarizing how long individual turns tend to be. bigram_entropy This feature measures the entropy of the bigram distribution, so higher values indicate greater diversity in adjacent token sequences. bigram_mutual_information This feature measures the average pointwise mutual information of adjacent token pairs, so higher values indicate that neighboring words are more strongly associated than expected by chance. burstiness_measure This feature measures the variability of utterance lengths relative to their mean, capturing how bursty or uneven the discourse is. character_entropy This feature measures the entropy of the non-whitespace character distribution, capturing low-level orthographic diversity. compression_ratio This feature measures the ratio of compressed to uncompressed text length, serving as a rough proxy for redundancy and repetitiveness in the corpus. conversations_lt_128_tokens pct This feature measures the proportion of conversations shorter than 128 tokens, capturing the prevalence of very short interactions. conversations_lt_512_tokens pct This feature measures the proportion of conversations shorter than 512 tokens, capturing how concentrated the dataset is in relatively short interactions. diminutive_frequency This feature summarizes the corpus property captured by diminutive frequency. dis_legomena_ratio This feature measures the proportion of vocabulary types that occur exactly twice in the corpus, providing a second view of lexical tail structure beyond hapax items. elongated_word_frequency This feature measures how often words with elongated character sequences appear, capturing expressive spellings such as stretched vowels or consonants. emotion_word_pct This feature measures the proportion of utterances containing words from an emotion-related lexicon, capturing how frequently affective vocabulary appears in the corpus. eventive_verb_pct This feature summarizes the corpus property captured by eventive verb pct. eventive_verb_token_pct This feature measures the proportion of tokens that belong to an eventive-verb lexicon, capturing how much the corpus emphasizes actions and events. frequency_distribution skewness This feature measures the skewness of the token-frequency distribution, capturing how strongly the corpus is dominated by a small number of very frequent words. hapax_ratio This feature measures the proportion of vocabulary types that occur exactly once in the corpus, capturing how much lexical mass lies in the long tail. interjection_frequency This feature measures how often interjections such as āohā or āwowā appear, capturing expressive and affective discourse style. intrinsic_perplexity_proxy This feature is a simple intrinsic predictability proxy derived from token entropy, with higher values indicating a less predictable token distribution. lexical_diversity_slope This feature measures how lexical diversity changes over time across the dataset, capturing whether later conversations become more or less varied in word choice. mattr This feature measures moving-average type-token ratio over a sliding window, giving a more length-robust estimate of lexical diversity than raw TTR. mean_utterance_length_slope This feature measures the linear change in average utterance length over the course of the dataset, capturing whether turns become longer or shorter over time. narrative_discourse_marker pct This feature measures the proportion of utterances containing narrative discourse markers such as temporal or sequencing words, capturing how often the corpus uses narrative-style structuring. narrative_discourse_marker token_pct This feature measures the proportion of tokens that are narrative discourse markers, capturing the token-level prevalence of narrative-style sequencing language. num_conversations This feature counts the number of conversations in the dataset, providing a coarse measure of dataset coverage and interactional breadth. num_sentences This feature summarizes the corpus property captured by num sentences. num_tokens This feature counts the total number of tokens in the corpus, providing a direct measure of dataset size. num_types This feature counts the number of unique token types in the corpus, providing a basic measure of vocabulary breadth. num_utterances This feature counts the number of utterances in the corpus, providing a basic measure of conversational segmentation and turn volume. past_tense_verb_pct This feature measures the average per-utterance rate of past-tense verbs, capturing how often speakers refer to past events. past_tense_verb_token_pct This feature measures the proportion of tokens that are identified as past-tense verbs by a simple heuristic, capturing past-event language at the token level. positive_sentiment_word_pct This feature summarizes the corpus property captured by positive sentiment word pct. praise_word_pct This feature summarizes the corpus property captured by praise word pct. sound_effect_frequency This feature measures how often onomatopoeic or sound-effect words occur, capturing playful and imitative sound language. syntactic_complexity_slope over_time_proxy This feature measures the change in a proxy for syntactic complexity over time, operationalized from shifts in average utterance length across the dataset. token_entropy This feature measures the entropy of the unigram token distribution, with higher values indicating a more diverse and less concentrated vocabulary. top1000_token_coverage This feature measures the proportion of all tokens covered by the 1,000 most frequent word types, capturing how concentrated the corpus is in a relatively small core vocabulary. top100_token_proportion This feature measures the proportion of all tokens covered by the 100 most frequent word types, capturing lexical concentration among the most common words. trigram_entropy This feature measures the entropy of the trigram distribution, so higher values indicate greater diversity and less predictability in local three-token sequences. ttr This feature measures the raw type-token ratio, or the number of unique tokens divided by total tokens, as a simple index of lexical diversity. unigram_entropy This feature measures the entropy of the unigram token distribution, with higher values indicating a more diverse and less predictable vocabulary. vocabulary_growth_over_time slope This feature measures the rate at which new vocabulary accumulates over the course of the dataset, capturing how quickly lexical novelty is introduced. zipf_slope This feature measures the slope of the log-rank versus log-frequency relationship, capturing how closely the vocabulary frequency distribution follows a Zipf-like pattern. POS / Syntactic Composition dep_avg_clause_count_per sentence This feature measures the average number of clause-related dependency relations per sentence, capturing how clause-dense the parsed language is. dep_avg_clause_dependency count_per_sentence_proxy This feature is a proxy for clause density based on clause-related dependency relations per sentence. dep_avg_dependency_length This feature measures the average linear distance between tokens and their syntactic heads, a standard proxy for syntactic locality or complexity. dep_avg_parse_tree_depth This feature measures average dependency parse-tree depth, which serves as a structural proxy for syntactic complexity. dep_avg_parse_tree_depth proxy This feature is a proxy version of average parse-tree depth used for analyses that rely on a compact syntactic-complexity summary. dep_dependency_length variance This feature measures the variance of dependency lengths, capturing how variable syntactic attachment distances are across the corpus. dep_left_branching_ratio This feature measures the proportion of dependency links that branch to the left, capturing one aspect of syntactic branching direction. dep_num_docs_parsed This feature counts how many utterance documents were successfully parsed, providing a measure of usable input for the dependency parser. dep_num_sentences This feature counts the number of parsed sentences, providing the sample size for sentence-level dependency metrics. dep_pct_coordinate_clauses This feature measures the proportion of parsed sentences containing coordination, capturing how often clauses are linked side by side. dep_pct_fragments_parse This feature measures the proportion of parsed sentences classified as fragments, capturing incomplete or very short syntactic units. dep_pct_imperatives_parse This feature measures the proportion of parsed sentences classified as imperatives, capturing directive clause structure. dep_pct_passive constructions This feature measures the proportion of parsed sentences containing passive markers, capturing how often passive syntax appears. dep_pct_questions_parse This feature measures the proportion of parsed sentences identified as questions, using the dependency-parse pipeline rather than simple punctuation alone. dep_pct_subordinate_clauses This feature measures the proportion of parsed sentences containing subordinate clause relations, capturing hierarchical syntactic embedding. dep_pct_wh_questions_parse This feature measures the proportion of parsed sentences that are WH-questions, capturing interrogatives introduced by words such as āwhatā or āwhereā. dep_pct_yes_no_questions parse This feature measures the proportion of parsed sentences that are yes/no questions, capturing interrogatives formed without WH-words. dep_right_branching_ratio This feature measures the proportion of dependency links that branch to the right, capturing one aspect of syntactic branching direction. dep_total_utterances available This feature counts the total number of utterances that were available before dependency parsing, providing a measure of syntactic evidence volume. dep_utterances_used_for parsing This feature counts how many utterances were actually passed to the dependency parser after any sampling or filtering. pos_num_tagged_tokens This feature counts the total number of POS-tagged tokens across the corpus, providing the amount of text included in the POS analysis pipeline. pos_num_tagged_tokens caregiver This feature counts the number of POS-tagged tokens produced by the caregiver, capturing how much caregiver speech is available for part-of-speech analysis. pos_num_tagged_tokens_child This feature counts the number of POS-tagged tokens produced by the child, capturing how much child speech is available for part-of-speech analysis. pos_pct_adjectives This feature measures the proportion of POS-tagged tokens that are adjectives, capturing how much descriptive modifier language appears. pos_pct_adverbs This feature measures the proportion of POS-tagged tokens that are adverbs, capturing manner, degree, and discourse-modifying language. pos_pct_conjunctions This feature measures the proportion of POS-tagged tokens that are conjunctions, capturing how often clauses or phrases are linked together. pos_pct_content_words This feature measures the proportion of POS-tagged tokens that are content words, capturing how much of the corpus consists of semantically rich lexical items. pos_pct_determiners This feature measures the proportion of POS-tagged tokens that are determiners, capturing the use of noun-phrase marking words such as articles and demonstratives. pos_pct_function_words This feature measures the proportion of POS-tagged tokens that are function words, capturing closed-class grammatical material. pos_pct_nouns This feature measures the proportion of POS-tagged tokens that are nouns, capturing how strongly the corpus emphasizes object and entity reference. pos_pct_nouns_caregiver This feature measures the proportion of caregiver POS-tagged tokens that are nouns, capturing noun usage specifically in caregiver speech. pos_pct_nouns_child This feature measures the proportion of child POS-tagged tokens that are nouns, capturing noun usage specifically in child speech. pos_pct_prepositions This feature measures the proportion of POS-tagged tokens that are prepositions or adpositions, capturing relational language. pos_pct_pronouns This feature measures the proportion of POS-tagged tokens that are pronouns, capturing referential language that depends on discourse context. pos_pct_verbs This feature measures the proportion of POS-tagged tokens that are verbs, capturing the amount of event- and action-centered language. pos_pct_verbs_caregiver This feature measures the proportion of caregiver POS-tagged tokens that are verbs, capturing verb usage specifically in caregiver speech. pos_pct_verbs_child This feature measures the proportion of child POS-tagged tokens that are verbs, capturing verb usage specifically in child speech. pos_pos_bigram_diversity This feature measures the diversity of observed POS bigrams relative to the number of POS transitions, capturing local syntactic variety. pos_pos_bigram_diversity tagspace This feature measures POS-bigram diversity normalized by the size of the possible tag space, capturing how much of the available local syntactic space is used. pos_pos_bigram_entropy This feature measures the entropy of adjacent POS-tag bigrams, capturing the diversity of local syntactic patterns in the corpus. pos_pos_entropy This feature measures entropy over POS tags, capturing how evenly different part-of-speech categories are distributed. pos_pos_entropy_caregiver This feature measures POS-tag entropy in caregiver speech, capturing syntactic-category diversity specifically for caregiver language. pos_pos_entropy_child This feature measures POS-tag entropy in child speech, capturing syntactic-category diversity specifically for child language. pos_total_utterances available This feature counts the total number of utterances available before the POS-tagging pipeline, providing the available sample size for POS analysis. pos_utterances_used_for_pos This feature counts how many utterances were actually used for POS tagging after any sampling or filtering. Conversational / Discourse Structure adjacent_turn_lexical overlap This feature measures lexical overlap between adjacent turns, capturing how much consecutive speakers reuse the same words. anaphora_usage_pct This feature measures the proportion of tokens that are third-person anaphoric forms such as pronouns, capturing discourse-dependent reference. avg_caregiver_utterance_len This feature measures the average length of caregiver utterances in tokens, summarizing how long caregiver turns tend to be. avg_child_utterance_len This feature measures the average length of child utterances in tokens, summarizing how long child turns tend to be. avg_conversation_len_turns This feature measures the average number of turns per conversation, capturing interactional depth independent of token length. caregiver_after_child_any overlap_pct This feature measures the proportion of child-to-caregiver adjacent pairs in which the caregiver reuses at least some child vocabulary. caregiver_child_token_ratio This feature measures the ratio of caregiver tokens to child tokens, summarizing the relative contribution of the two speaker groups. caregiver_question_pct This feature measures the proportion of caregiver utterances that are questions, capturing how often caregivers prompt or query the child. caregiver_reuse_child_tokens pct This feature measures the proportion of child-to-caregiver adjacent pairs in which the caregiver reuses child vocabulary, operationalized as lexical overlap greater than zero. caregiver_token_pct This feature measures the proportion of all corpus tokens produced by caregivers. caregiver_wh_question_pct This feature measures the proportion of caregiver utterances that are WH-questions, capturing a more specific type of information-seeking prompt. child_token_pct This feature measures the proportion of all corpus tokens produced by the child. clarification_pct This feature measures the proportion of utterances containing clarification markers or repair-oriented questions, capturing requests for clearer information. confirmation_pct This feature measures the proportion of utterances containing confirmation markers such as āyesā or āokayā, capturing explicit acknowledgement or agreement. directive_pct This feature measures the proportion of utterances classified as directives or commands, capturing instruction-like language. expansion_pct This feature measures the proportion of utterances that contain expansion markers suggesting elaboration or reformulation of prior content. fragment_pct This feature measures the proportion of utterances classified as fragments, typically because they are very short or incomplete. imperative_pct This feature measures the proportion of utterances classified as imperatives, capturing command-like speech. overlapping_speech_marker pct This feature measures the rate of overlap markers in the transcript, capturing instances where speech overlaps or is explicitly annotated as overlapping. pronoun_density This feature measures the proportion of all tokens that are pronouns, capturing the density of discourse-dependent reference. question_pct This feature measures the proportion of utterances that are questions, based primarily on question-mark cues. self_repair_frequency This feature measures how often utterances contain fillers or repair markers such as restarts, corrections, or hesitation expressions. speaker_label_consistency This feature measures the proportion of utterances whose speaker labels are recognized and consistent with the expected label set. topic_entropy_overlap heuristic_proxy This feature is an overlap-based proxy for topical variation, using a coarse heuristic intended to capture topic dispersion rather than a full topic model. topic_persistence_length overlap_heuristic_proxy This feature measures the average run length of local lexical-overlap segments, capturing how long a topic tends to persist across turns. turn_taking_frequency This feature measures how often speaker identity switches across adjacent turns, capturing the pace of turn alternation. unknown_speaker_utterance pct This feature measures the proportion of utterances with unknown or unrecognized speaker labels. wh_question_pct This feature measures the proportion of utterances that are WH-questions, capturing interrogatives introduced by question words. yes_no_question_pct This feature measures the proportion of utterances that are yes/no questions rather than WH-questions. Semantic Coherence sem_mean_adjacent_turn semantic_similarity This feature measures the average embedding-based semantic similarity between adjacent turns, capturing how semantically connected neighboring utterances are. sem_mean_adjacent_turn semantic_similarity_len weighted This feature measures the average semantic similarity between adjacent turns while weighting pairs by turn length, so longer turns contribute more to the summary. sem_mean_semantic_similarity caregiver_to_child This feature measures the average semantic similarity from caregiver turns to immediately following child turns, capturing one direction of interactional semantic alignment. sem_mean_semantic_similarity child_to_caregiver This feature measures the average semantic similarity from child turns to immediately following caregiver turns, capturing the reverse direction of interactional semantic alignment. sem_num_adjacent_pairs_used This feature counts the number of adjacent utterance pairs included in the semantic-similarity computation after filtering and sampling. sem_num_pairs_caregiver_to child This feature counts the number of caregiver-to-child adjacent turn pairs available for directional semantic-similarity analysis. sem_num_pairs_child_to caregiver This feature counts the number of child-to-caregiver adjacent turn pairs available for directional semantic-similarity analysis. sem_total_adjacent_pairs available This feature counts all adjacent utterance pairs available before semantic-similarity sampling or filtering. sem_total_utterances available This feature counts the total number of utterances available before the semantic-similarity pipeline is applied. sem_utterances_used_for semantic This feature counts how many utterances were actually used when constructing adjacent pairs for semantic analysis. Repetition / Formulaicity exact_sentence_repetition rate This feature measures how often whole utterances repeat exactly, capturing formulaic reuse at the sentence level. ngram_repetition_rate This feature measures the overall rate of repeated bigrams and trigrams, capturing formulaicity and local repetition in the corpus. partial_repetition_frequency This feature measures how often utterances contain partial repetition or reduplication-like patterns, capturing local repetition within turns. reduplication_frequency This feature measures how often explicit reduplication patterns occur, such as repeated words or hyphenated duplicates. repeated_bigram_ratio This feature measures the proportion of all bigram tokens that belong to repeated bigram types, capturing repeated two-word sequences. repeated_trigram_ratio This feature measures the proportion of all trigram tokens that belong to repeated trigram types, capturing repeated three-word sequences. Mixture / Composition Metadata age_age_metadata_alignment score This feature summarizes how tightly concentrated the ages of the component families are within a mixture, with higher values indicating stronger age alignment. age_num_component_families This feature counts how many component families contribute to a mixture in the age-metadata analysis. age_num_families_with_age This feature counts how many component families in a mixture have usable age metadata available. age_weighted_age_std_months This feature measures the weighted standard deviation of component-family ages in months, capturing age heterogeneity within a mixture. age_weighted_mean_age_months This feature measures the weighted mean age in months of component families contributing to a mixture. div_max_js_with_others This feature measures the maximum JensenāShannon divergence between a dataset and any other dataset, capturing its strongest distributional distinctiveness. div_mean_js_with_others This feature measures the mean JensenāShannon divergence between a dataset and all other datasets, capturing average distributional distinctiveness. div_mean_kl_from_others This feature measures the mean KL divergence from other datasets into the current one, capturing how distinct the dataset is when others are compared against it. div_mean_kl_to_others This feature measures the mean KL divergence from the current dataset to other datasets, capturing how distinct its token distribution is relative to the rest. div_median_js_with_others This feature measures the median JensenāShannon divergence between a dataset and all others, providing a robust summary of distributional distinctiveness. mix_content_vocab_support size This feature counts the size of the support of content-word vocabulary in the mixture distribution after formatting and non-content tokens are removed. mix_cross_family_lexical overlap This feature measures average lexical overlap across component families in a mixture, capturing how similar their vocabularies are. mix_effective_vocab_size bits This feature converts mixture token entropy in bits into an effective vocabulary size, summarizing how broadly probability mass is spread across word types. mix_effective_vocab_size content_only This feature is the effective vocabulary size computed after removing formatting-style tokens, capturing lexical breadth in content words only. mix_effective_vocab_size nats This feature converts mixture token entropy in nats into an effective vocabulary size, providing an alternative entropy-scale summary of lexical breadth. mix_effective_vocabulary size This feature summarizes the effective number of vocabulary items represented in the mixture distribution based on entropy. mix_family_diversity_index This feature measures how evenly a mixture is distributed across its component families, with higher values indicating a more balanced mixture. mix_formatting_token_mass This feature measures how much probability mass in the mixture distribution is assigned to formatting or metadata-like tokens. mix_inter_family_variance entropy This feature measures variance in per-family token entropy within a mixture, capturing heterogeneity in lexical concentration across families. mix_inter_family_variance syntactic_complexity This feature measures variance in a syntactic-complexity proxy across families within a mixture, capturing cross-family heterogeneity in utterance structure. mix_js_divergence_within mixture This feature measures average pairwise JensenāShannon divergence among component families within a mixture, capturing internal distributional heterogeneity. mix_mixture_token_entropy bits This feature measures the entropy of the mixture token distribution in bits, capturing how diverse and even the overall mixture vocabulary is. mix_mixture_token_entropy nats This feature measures the entropy of the mixture token distribution in natural-log units, providing the same information on a different scale. mix_mixture_top100_token mass This feature measures the fraction of mixture probability mass captured by the 100 most probable tokens, reflecting lexical concentration. mix_mixture_top10_token_mass This feature measures the fraction of mixture probability mass captured by the 10 most probable tokens, reflecting lexical concentration in a small core vocabulary. mix_mixture_top1_token_mass This feature measures the fraction of mixture probability mass assigned to the single most probable token. mix_mixture_top50_token_mass This feature measures the fraction of mixture probability mass captured by the 50 most probable tokens. mix_mixture_top5_token_mass This feature measures the fraction of mixture probability mass captured by the 5 most probable tokens. mix_mixture_vocab_support size This feature counts how many token types have nonzero probability in the overall mixture distribution. mix_mixture_weight_entropy bits This feature measures the entropy of the component-family mixture weights, capturing how evenly the mixture draws from its sources. mix_num_component_families This feature counts the number of component families contributing nonzero weight to a mixture. mix_token_entropy_bits content_only This feature measures token entropy in bits after removing formatting-like tokens, capturing content-word diversity. mix_vocab_support_size This feature counts how many token types have nonzero probability in the mixture and functions as a support-size summary of vocabulary breadth. Reference LM Predictability ppl_norm_babylm20m_caregiver only This feature measures normalized perplexity under a BabyLM-trained reference language model on caregiver-only text, capturing how predictable the caregiver stream is to that model. ppl_norm_babylm20m_full dialogue This feature measures normalized perplexity under a BabyLM-trained reference language model on the full dialogue stream, capturing overall corpus predictability. ppl_norm_babylm20m_mean scopes This feature averages the BabyLM reference-model perplexity summaries across scoring scopes, providing a broader predictability estimate. ppl_norm_external_gpt2 caregiver_only This feature measures normalized perplexity under an external GPT-2 model on caregiver-only text, capturing how predictable caregiver language is to a general-domain model. ppl_norm_external_gpt2_full dialogue This feature measures normalized perplexity under an external GPT-2 model on the full dialogue stream, capturing overall predictability to a general-domain model. ppl_norm_external_gpt2_mean scopes This feature averages the external GPT-2 perplexity summaries across scoring scopes, providing a broader predictability estimate. Transcription / Data Quality partial_word_pct This feature measures the rate of partial-word markers in the transcript, capturing truncations or cut-off productions. unintelligible_marker_pct This feature measures the rate of unintelligible markers such as āxā or related annotations, capturing transcription uncertainty or inaudible speech. Table 9: Full linguistic feature results grouped by category, sorted by Top10 hit count within each category. Top10 is the total number of modelātarget analyses in which a feature appeared in a methodās top-10 list; Meth. is the number of methods (Spearman, lasso, XGBoost) for which the feature appeared at least once; Spear., Lasso, and XGB give the corresponding method-specific hit counts; Mean rank is the average within-method top-10 rank across all hits; Mean |Ļ||Ļ| is the mean absolute Spearman correlation across the featureās Spearman top-10 appearances; Mean |β||β| is the mean absolute lasso coefficient across the featureās lasso top-10 appearances; Mean imp. is the mean XGBoost feature-importance score across the featureās XGBoost top-10 appearances. Feature Top10 Meth. Spear. Lasso XGB Mean rank Mean |Ļ||Ļ| Mean |β||β| Mean imp. Lexical / Distributional / Scale bigram_mutual_information 9 3 3 1 5 5.56 0.72 0.153 0.060 hapax_ratio 5 3 1 1 3 6.60 0.48 0.264 0.067 trigram_entropy 5 2 4 0 1 5.00 0.74 0.000 0.075 num_conversations 5 1 0 5 0 2.00 0.00 0.178 0.000 avg_conversation_len_tokens 4 3 2 1 1 8.25 0.57 0.005 0.036 bigram_entropy 4 2 1 0 3 4.25 0.87 0.000 0.065 emotion_word_pct 4 2 1 3 0 5.75 0.40 0.013 0.000 mean_utterance_length_slope 4 2 1 0 3 6.00 0.50 0.000 0.052 conversations_lt 512_tokens_pct 3 3 1 1 1 7.67 0.56 0.331 0.020 num_types 3 2 2 1 0 2.33 0.82 0.014 0.000 narrative_discourse marker_pct 3 2 1 0 2 4.33 0.48 0.000 0.044 syntactic_complexity slope_over_time_proxy 3 2 1 0 2 6.67 0.50 0.000 0.054 burstiness_measure 3 2 0 1 2 6.67 0.00 0.218 0.039 narrative_discourse marker_token_pct 3 2 2 0 1 7.33 0.43 0.000 0.014 frequency distribution_skewness 3 1 3 0 0 4.67 0.72 0.000 0.000 intrinsic_perplexity_proxy 3 1 0 0 3 7.00 0.00 0.000 0.037 dis_legomena_ratio 2 2 0 1 1 6.00 0.00 0.004 0.034 elongated_word_frequency 2 2 0 1 1 7.50 0.00 0.002 0.031 eventive_verb_token_pct 2 2 0 1 1 8.00 0.00 0.001 0.029 num_tokens 2 2 1 0 1 8.50 0.84 0.000 0.022 lexical_diversity_slope 2 1 0 0 2 5.50 0.00 0.000 0.041 sound_effect_frequency 1 1 0 0 1 2.00 0.00 0.000 0.108 conversations_lt 128_tokens_pct 1 1 0 1 0 4.00 0.00 0.002 0.000 past_tense_verb_pct 1 1 0 1 0 4.00 0.00 0.350 0.000 num_utterances 1 1 1 0 0 7.00 0.85 0.000 0.000 mattr 1 1 0 0 1 7.00 0.00 0.000 0.036 vocabulary_growth over_time_slope 1 1 0 0 1 7.00 0.00 0.000 0.039 compression_ratio 1 1 1 0 0 9.00 0.38 0.000 0.000 past_tense_verb_token_pct 1 1 0 0 1 9.00 0.00 0.000 0.035 interjection_frequency 1 1 0 0 1 10.00 0.00 0.000 0.032 avg_sentence_len_tokens 0 0 0 0 0 ā 0.00 0.000 0.000 avg_utterance_len_tokens 0 0 0 0 0 ā 0.00 0.000 0.000 character_entropy 0 0 0 0 0 ā 0.00 0.000 0.000 diminutive_frequency 0 0 0 0 0 ā 0.00 0.000 0.000 eventive_verb_pct 0 0 0 0 0 ā 0.00 0.000 0.000 num_sentences 0 0 0 0 0 ā 0.00 0.000 0.000 positive_sentiment_word_pct 0 0 0 0 0 ā 0.00 0.000 0.000 praise_word_pct 0 0 0 0 0 ā 0.00 0.000 0.000 token_entropy 0 0 0 0 0 ā 0.00 0.000 0.000 top1000_token_coverage 0 0 0 0 0 ā 0.00 0.000 0.000 top100_token_proportion 0 0 0 0 0 ā 0.00 0.000 0.000 ttr 0 0 0 0 0 ā 0.00 0.000 0.000 unigram_entropy 0 0 0 0 0 ā 0.00 0.000 0.000 zipf_slope 0 0 0 0 0 ā 0.00 0.000 0.000 POS / Syntactic Composition pos_num_tagged_tokens_child 7 3 1 4 2 2.00 0.86 0.604 0.356 pos_num_tagged tokens_caregiver 7 3 4 1 2 4.29 0.74 0.016 0.062 pos_pos_bigram_entropy 6 3 1 1 4 3.33 0.56 0.263 0.069 dep_total utterances_available 5 3 1 1 3 3.40 0.85 0.012 0.149 pos_pos_entropy_caregiver 4 3 1 1 2 4.25 0.64 0.003 0.061 dep_num_sentences 4 2 3 1 0 4.75 0.70 0.530 0.000 pos_pos_bigram diversity_tagspace 3 3 1 1 1 4.67 0.28 0.003 0.121 pos_pct_nouns_caregiver 3 2 1 0 2 2.67 0.37 0.000 0.059 dep_pct_coordinate_clauses 3 2 1 0 2 4.00 0.40 0.000 0.066 pos_pct_adverbs 3 2 1 0 2 5.00 0.45 0.000 0.029 pos_pos_entropy 3 2 2 0 1 5.33 0.49 0.000 0.073 dep_num_docs_parsed 3 2 1 0 2 7.33 0.49 0.000 0.027 pos_pos_bigram_diversity 3 1 3 0 0 9.33 0.69 0.000 0.000 dep_avg_parse_tree_depth 2 2 1 0 1 1.50 0.57 0.000 0.085 pos_pos_entropy_child 2 2 1 1 0 2.50 0.60 0.440 0.000 dep_avg_parse tree_depth_proxy 2 2 1 0 1 3.00 0.57 0.000 0.058 dep_right_branching_ratio 2 2 1 0 1 4.00 0.44 0.000 0.025 pos_pct_conjunctions 2 2 1 0 1 4.00 0.42 0.000 0.055 dep_pct_wh_questions_parse 2 2 1 1 0 7.50 0.32 0.001 0.000 pos_num_tagged_tokens 2 2 1 0 1 8.50 0.84 0.000 0.023 pos_pct_nouns_child 2 2 1 0 1 8.50 0.31 0.000 0.041 pos_pct_content_words 2 1 0 0 2 7.00 0.00 0.000 0.047 pos_pct_prepositions 2 1 0 2 0 7.00 0.00 0.268 0.000 dep_utterances used_for_parsing 1 1 1 0 0 2.00 0.86 0.000 0.000 dep_pct_passive constructions 1 1 1 0 0 3.00 0.36 0.000 0.000 dep_avg_clause count_per_sentence 1 1 1 0 0 4.00 0.53 0.000 0.000 dep_avg_clause_dependency count_per_sentence_proxy 1 1 1 0 0 5.00 0.53 0.000 0.000 dep_left_branching_ratio 1 1 0 1 0 5.00 0.00 0.002 0.000 pos_pct_verbs_caregiver 1 1 0 1 0 7.00 0.00 0.001 0.000 pos_total utterances_available 1 1 1 0 0 8.00 0.85 0.000 0.000 pos_pct_determiners 1 1 0 0 1 8.00 0.00 0.000 0.032 pos_pct_verbs 1 1 0 1 0 8.00 0.00 0.348 0.000 pos_pct_verbs_child 1 1 0 0 1 8.00 0.00 0.000 0.011 dep_pct_questions_parse 1 1 1 0 0 9.00 0.31 0.000 0.000 pos_utterances_used_for_pos 1 1 1 0 0 9.00 0.85 0.000 0.000 dep_avg_dependency_length 1 1 0 0 1 9.00 0.00 0.000 0.034 pos_pct_pronouns 1 1 1 0 0 10.00 0.27 0.000 0.000 dep_pct_fragments_parse 1 1 0 0 1 10.00 0.00 0.000 0.016 dep_dependency length_variance 0 0 0 0 0 ā 0.00 0.000 0.000 dep_pct_imperatives_parse 0 0 0 0 0 ā 0.00 0.000 0.000 dep_pct_subordinate_clauses 0 0 0 0 0 ā 0.00 0.000 0.000 dep_pct_yes_no questions_parse 0 0 0 0 0 ā 0.00 0.000 0.000 pos_pct_adjectives 0 0 0 0 0 ā 0.00 0.000 0.000 pos_pct_function_words 0 0 0 0 0 ā 0.00 0.000 0.000 pos_pct_nouns 0 0 0 0 0 ā 0.00 0.000 0.000 Semantic Coherence sem_num_pairs child_to_caregiver 6 3 2 2 2 6.17 0.80 0.005 0.050 sem_num_pairs caregiver_to_child 4 3 2 1 1 3.75 0.85 0.241 0.041 sem_total_adjacent pairs_available 4 3 2 1 1 5.25 0.80 0.762 0.090 sem_mean_adjacent_turn semantic_similarity 3 3 1 1 1 5.67 0.43 0.001 0.044 sem_num_adjacent_pairs_used 3 2 2 1 0 6.00 0.79 0.008 0.000 sem_mean_adjacent_turn_semantic similarity_len_weighted 1 1 1 0 0 5.00 0.41 0.000 0.000 sem_mean_semantic_similarity caregiver_to_child 1 1 1 0 0 7.00 0.51 0.000 0.000 sem_mean_semantic_similarity child_to_caregiver 1 1 1 0 0 8.00 0.39 0.000 0.000 sem_total utterances_available 1 1 1 0 0 10.00 0.85 0.000 0.000 sem_utterances used_for_semantic 0 0 0 0 0 ā 0.00 0.000 0.000 Mixture / Composition Metadata div_mean_kl_from_others 5 3 3 1 1 3.60 0.71 0.500 0.053 mix_inter_family variance_entropy 4 3 1 1 2 5.50 0.60 0.074 0.056 mix_cross_family lexical_overlap 4 2 3 0 1 5.50 0.49 0.000 0.078 mix_family_diversity_index 4 2 2 2 0 5.50 0.42 0.013 0.000 div_median_js_with_others 3 3 1 1 1 3.33 0.86 0.536 0.041 age_weighted_mean_age_months 3 2 2 1 0 6.33 0.31 0.013 0.000 div_mean_js_with_others 3 1 0 0 3 4.00 0.00 0.000 0.080 mix_content_vocab support_size 2 2 1 0 1 1.50 0.85 0.000 0.332 mix_mixture top100_token_mass 2 2 1 0 1 1.50 0.64 0.000 0.170 mix_effective vocab_size_bits 2 2 1 0 1 2.00 0.61 0.000 0.236 div_mean_kl_to_others 2 2 1 0 1 3.50 0.42 0.000 0.037 age_num_component_families 2 2 1 0 1 5.00 0.54 0.000 0.058 mix_mixture_top10_token_mass 1 1 0 0 1 1.00 0.00 0.000 0.106 mix_mixture_top50_token_mass 1 1 1 0 0 2.00 0.62 0.000 0.000 mix_mixture_vocab support_size 1 1 1 0 0 3.00 0.85 0.000 0.000 mix_effective vocab_size_nats 1 1 1 0 0 4.00 0.61 0.000 0.000 mix_vocab_support_size 1 1 1 0 0 4.00 0.85 0.000 0.000 mix_effective vocabulary_size 1 1 1 0 0 5.00 0.61 0.000 0.000 mix_inter_family_variance syntactic_complexity 1 1 0 0 1 5.00 0.00 0.000 0.045 mix_mixture_token entropy_bits 1 1 1 0 0 6.00 0.61 0.000 0.000 mix_formatting_token_mass 1 1 0 0 1 6.00 0.00 0.000 0.051 mix_mixture_token entropy_nats 1 1 1 0 0 7.00 0.61 0.000 0.000 mix_mixture_top5_token_mass 1 1 0 0 1 7.00 0.00 0.000 0.034 age_age_metadata alignment_score 1 1 0 0 1 8.00 0.00 0.000 0.039 div_max_js_with_others 1 1 0 0 1 8.00 0.00 0.000 0.035 age_num_families_with_age 1 1 1 0 0 9.00 0.54 0.000 0.000 mix_num_component_families 1 1 1 0 0 10.00 0.54 0.000 0.000 age_weighted_age_std_months 0 0 0 0 0 ā 0.00 0.000 0.000 mix_effective_vocab size_content_only 0 0 0 0 0 ā 0.00 0.000 0.000 mix_js_divergence within_mixture 0 0 0 0 0 ā 0.00 0.000 0.000 mix_mixture_top1_token_mass 0 0 0 0 0 ā 0.00 0.000 0.000 mix_mixture_weight entropy_bits 0 0 0 0 0 ā 0.00 0.000 0.000 mix_token_entropy bits_content_only 0 0 0 0 0 ā 0.00 0.000 0.000 Conversational / Discourse Structure topic_persistence_length overlap_heuristic_proxy 4 3 1 2 1 3.75 0.51 0.002 0.059 avg_conversation_len_turns 3 3 1 1 1 4.00 0.59 0.432 0.077 speaker_label_consistency 3 3 1 1 1 5.00 0.59 0.003 0.019 self_repair_frequency 3 3 1 1 1 5.33 0.54 0.353 0.066 overlapping speech_marker_pct 3 2 1 2 0 5.00 0.50 0.015 0.000 child_token_pct 3 2 0 2 1 6.67 0.00 0.184 0.019 fragment_pct 2 2 1 0 1 6.00 0.54 0.000 0.033 anaphora_usage_pct 2 2 1 0 1 7.00 0.53 0.000 0.035 turn_taking_frequency 2 1 0 0 2 7.00 0.00 0.000 0.039 caregiver_child_token_ratio 1 1 0 1 0 1.00 0.00 0.838 0.000 unknown_speaker utterance_pct 1 1 1 0 0 2.00 0.59 0.000 0.000 directive_pct 1 1 0 1 0 2.00 0.00 0.004 0.000 caregiver_token_pct 1 1 0 1 0 3.00 0.00 0.004 0.000 wh_question_pct 1 1 1 0 0 4.00 0.36 0.000 0.000 caregiver_question_pct 1 1 1 0 0 6.00 0.34 0.000 0.000 expansion_pct 1 1 0 1 0 6.00 0.00 0.090 0.000 question_pct 1 1 1 0 0 7.00 0.34 0.000 0.000 caregiver_after child_any_overlap_pct 1 1 1 0 0 8.00 0.49 0.000 0.000 caregiver_reuse child_tokens_pct 1 1 1 0 0 9.00 0.49 0.000 0.000 pronoun_density 1 1 1 0 0 9.00 0.27 0.000 0.000 imperative_pct 1 1 0 1 0 9.00 0.00 0.000 0.000 avg_child_utterance_len 1 1 1 0 0 10.00 0.49 0.000 0.000 adjacent_turn lexical_overlap 1 1 0 1 0 10.00 0.00 0.000 0.000 avg_caregiver_utterance_len 0 0 0 0 0 ā 0.00 0.000 0.000 caregiver_wh_question_pct 0 0 0 0 0 ā 0.00 0.000 0.000 clarification_pct 0 0 0 0 0 ā 0.00 0.000 0.000 confirmation_pct 0 0 0 0 0 ā 0.00 0.000 0.000 topic_entropy_overlap heuristic_proxy 0 0 0 0 0 ā 0.00 0.000 0.000 yes_no_question_pct 0 0 0 0 0 ā 0.00 0.000 0.000 Repetition / Formulaicity repeated_trigram_ratio 4 3 1 1 2 5.25 0.37 0.004 0.085 ngram_repetition_rate 4 2 1 0 3 2.75 0.31 0.000 0.116 exact_sentence repetition_rate 2 2 1 0 1 4.00 0.50 0.000 0.145 repeated_bigram_ratio 2 1 0 0 2 5.00 0.00 0.000 0.118 partial_repetition_frequency 1 1 0 1 0 8.00 0.00 0.011 0.000 reduplication_frequency 0 0 0 0 0 ā 0.00 0.000 0.000 Reference LM Predictability ppl_norm_babylm20m full_dialogue 3 2 0 1 2 6.33 0.00 0.012 0.050 ppl_norm_babylm20m caregiver_only 2 1 0 2 0 6.50 0.00 0.003 0.000 ppl_norm_babylm20m mean_scopes 2 1 0 0 2 9.50 0.00 0.000 0.026 ppl_norm_external gpt2_mean_scopes 1 1 0 0 1 10.00 0.00 0.000 0.019 ppl_norm_external gpt2_caregiver_only 0 0 0 0 0 ā 0.00 0.000 0.000 ppl_norm_external gpt2_full_dialogue 0 0 0 0 0 ā 0.00 0.000 0.000 Transcription / Data Quality unintelligible_marker_pct 2 2 0 1 1 9.00 0.00 0.001 0.031 partial_word_pct 0 0 0 0 0 ā 0.00 0.000 0.000