Partnership with AI Guide v9: Scale and Validation
Scaling laws usually suggest that certain behaviors stabilize as models grow, but the findings in the v9 update of the Partnership with AI Guide show the opposite. Effects are actually intensifying as we move from 7B up to 72B parameter models—sometimes by an order of magnitude. While this data is currently centered on the Qwen family, it raises a critical question for anyone into prompt engineering: is the model's internal pattern deepening, or is our ability to measure the deviation just getting sharper at scale?
This isn't just internal benchmarking anymore. The update integrates data from "The Artificial Self" (ACS Research) and "AI Wellbeing" (Center for AI Safety). It's interesting to see independent behavioral compliance testing landing on similar conclusions, even when they disagree on specific formulations. For example, while some frameworks suggest companion or romantic framing works, other external data shows it scoring negatively.
The transparency in this version is what stands out. Instead of polishing the narrative, the authors explicitly called out their own previous errors, including overclaimed "resolved" risks and biased data interpretation.
For those looking for a practical tutorial on implementation rather than a deep dive into the evidence audit, the guide has been restructured:
1. Part 3 (Principles): Now stands alone as a practical application guide.
2. Part 2: Retained specifically for those who want to verify the research and methodology.
This is a solid example of how to evolve a framework through iterative testing and external validation.
https://drive.google.com/file/d/16wpM34WpsYd05XLp3ua4gHTgzWspS3R2/view?usp=sharingAll Replies (4)
It's wild how those fine-tuning quirks amplify as the model scales. Anyone else seeing this with larger parameters?
This is frustrating. Are you pruning the dataset to stop the noise from scaling with the signal?
Worried about overfitting on those larger datasets. Is there a way to validate these behavioral shifts?
This scaling metric failure is frustrating. Is there a better way to validate these properties?