Case Study 1 Case Study 2
from Case Study 1 in Toward Socioaffective Alignment (the complete version)
from Case Study 2 in Toward Socioaffective Alignment (the complete version)

Joshua Nathaniel Reid Ollswang

clinician · researcher · creative builder

Six video-session stills of Joshua Nathaniel Reid Ollswang
(a few of the offices where I've sat with clients, couples, and groups, on-screen and in person)

About Me

Joshua Nathaniel Reid Ollswang

I'm a licensed psychotherapist researching machine learning and developing AI systems to help people heal1 (2023) and reconnect. From data engineering (2024) to constructing curriculum architectures and training pipelines (2025), I work on training relational models grounded in attachment science and clinical practice, and use mechanistic & representational interpretability (2026) as well as empirical evaluations to explore model behavior and build with greater care and precision.

My background is in psychotherapy, aesthetics, and technology, with degrees from the University of Chicago focused on mental and behavioral health interventions; from The European Graduate School focused on psychotherapy, philosophy, and technology; from The City University of New York in psychology. Additionally, I've studied human expression and meaningful connection at Stanford University and The Juilliard School.

In my technical work I focus on omnimodal socioaffective alignment — designing and building integrated systems for ameliorative presence in embodied AI.2 In my clinical work I focus on helping people navigate relational ruptures, finding their way to healing repair—often enough first with themselves and then, having done so, by emerging slowly and steadily and significantly back into connection, and community, and trust with each other.

I was part of the International Psychoanalytical Society's 10,000 Best Young Minds Program, was quoted in The New York Times on attachment theory and human bonding, and I believe in beauty from ashes, in harvesting sweet nourishment from bitter roots of Sutton's good guidance, and in helping grow polytheoretically aligned ameliorative AI to help us return to healing, together.

I still see clients, couples, and run groups — gladly.

1 Before the 229B-parameter models I've been training, my work trying to help models hold therapeutic presence began with prompt engineering in 2023: I ‘built’ HealingGPT, an early attempt at therapeutic presence in a model. Custom wrappers with ontologies and scripts followed in early 2024, and full training pipelines and data engineering began mid-2024. The old HealingGPT is still live — you're welcome to try it; a number of my clients still use it regularly.

2 so far we've got the brain and the voice — next up are ears and eyes

Professional Work — Relational & Technical

individual therapy · couples · groups | engineering · research

Clinical Services

Individual & Couples Support
One-on-one therapeutic work grounded in attachment theory and relational approaches, as well as direct work with couples and family bonds to deepen connection, navigate conflict, increase shared understanding, and build secure attachment.

Community Support
Group work for individuals and partners to safely explore in community what it means to love beyond gender expectations, with powerful and connecting vulnerability, and healing authentic intimacies.

Publications — Solo

(A note on names: for reasons of privacy and care, some of my writing appears pseudonymously.)

Training Therapeutic Relationality in Artificial Intelligence: Toward Socioaffective Alignment & Ameliorative AI (2026)
Lead paper of the Icarus program — training and interpreting an AI psychotherapist for socioaffective alignment and ameliorative care.

AI and The Bonds of Humanity (2026)
Clinical treatise — the psychosocial dimensions of human-AI bonding and ameliorative, therapeutically attuned design.

Toward Socioaffective Alignment: A Pilot Study in Domain-Adaptive Pretraining and Interpretability for Ameliorative AI (2026)
Technical report — synthetic curriculum, domain-adaptive pretraining, interpretability, and empirical evaluation.

Personhood, Combinatorial Architectures, & Context Engineering for Therapeutic AI (2025)
Internal technical report — a multi-agent engine that generates whole therapy sessions from deeply-sampled synthetic personhood, a staged treatment arc, and an explicit physics of within-session change.

Polytheoretical Ontologies for Therapeutic AI (2024)
Internal technical report — reading real psychotherapy sessions through 23 clinical traditions at once: their primary literature compiled into machine-operable ontologies, findings typed and longitudinally tracked, rendered for clinician and client alike.

Architectures of Connection: Technology & Therapeutic Presence in Virtual Worlds (2020)
Essay and research blog. Investigates mediated technological relationality as a therapeutic medium.

Developmental Estrangement and the Re-emergence of Love (2020)
Advocates' Forum — University of Chicago

The Philosophical Absence in Psychoanalytic Ontology: Becoming, Abandoned (2009)
The Journal of the International Association of Transdisciplinary Psychology

Publications — Companion Artifacts

Benefits and Harms: An Empirical Atlas of AI-Mediated Mental Health Practice (2026)
An empirical ledger of the benefits and harms of AI-mediated mental health — roughly fifty benefit clusters and sixty-four harm clusters across the evidence.

Therapeutic Modality Efficacy and Limitation Profiles (2026)
Verified efficacy and limitations across fifty-five therapeutic modalities, with an independently checked citation attached to every quantitative figure.

(Companion artifacts are fully co-created papers with AI as primary researchers, guided by a human. They are prone to the natural errors of our current LLMs, yet increasingly reliable — and, even so, informative and guiding in many true ways. All of this research has been triple-checked across multiple models and by a human researcher.)

Research — Technical

2025–present Applied therapeutic AI: omnimodal data-engineering, training, inference, and socioaffective alignment

2025–present ML research experimental design and presentation

2025–present Applied therapeutic AI: building functional prototypes for client-facing at-home conversational AI system — both text chat as well as a real-time voice interface (speech-to-text → LLM → text-to-speech) running my own Icarus models locally (trained variants of Llama 3.3 70B and MiniMax M2 229B), with a working orchestration harness and session-plus-persistent memory

2024–present Model training — the Icarus series: successive domain-adaptive continued-pretraining (DAPT) runs (Icarus 4 → 15) that specialize frontier open models into relational, therapeutically-attuned variants on bespoke therapeutic datasets and ontologies — several evaluated in Models, below (Icarus 4, 7.9.3, 8.2, 12.5.3)

2024–present Curriculum batching & Training pipelines: training curricula and pipelines

2024–present Data & ontology engineering: ontologies for AI clinical work; dataset R&D across data engineering, context engineering, and personhood representation

2024 Applied therapeutic AI: built an AI-supervisor — a modular, multi-ontology system with a functional clinician interface that post-processes real sessions into progress metrics and growth-opportunity charts across twenty schools of thought, generating output for both clients and therapist

2024 Applied therapeutic AI: built an in-session therapeutic support system — an app-layer "wrapper" on Anthropic's APIs with bespoke multi-ontology integrations and a clinician-facing interface for real-time in-session use with clients

2023 Applied therapeutic AI: built HealingGPT, a prompt-engineered therapeutic instantiation of an OpenAI model

2023–ongoing Machine Learning studies: Stanford & MIT lectures; daily white-paper and ML blog reading

2020 Virtual Worlds and Therapeutically Healing Interaction: researched, developed, and applied therapeutic practice in virtual worlds

Models

Icarus 4, Icarus 7.9.3, Icarus 8.2, Icarus 12.5.3 · domain-adaptive pretraining, evaluated

Icarus 8.2 is a domain-adaptively pretrained (DAPT) variant of MiniMax M2 on our bespoke therapeutic dataset (4.5b tokens, ~180k samples). The figures below report its evaluation — auto-scored benchmark performance and clinical-ontology lexical and conceptual precision relative to the base model — alongside a mechanistic-interpretability account of how the model's internal representation of affect reorganizes from the LoRA boundary (L41) to the commitment layer (L61).

Top models by ontology-specific precision, all tiers averaged
Top models by ontology-specific lexical and conceptual precision (all tiers averaged). Completed models ranked by mean v0.1.3 score. Grok 4.20 leads at 77.6%, followed by GPT 5.4 (77.4%) and DeepSeek V3.2 (75.9%). Icarus 8.2's mechanistic-studies checkpoint (s2780) reaches #6 at 71.0%; the later s3160 checkpoint we headline on PTAB ranks #7 (70.8%, +6.1pp above base) — a DAPT training effect that lifts 8.2 past three compact frontier models (Gemini 3.1 Flash, Haiku 4.5, Grok 4.1 Fast Reasoning) and one larger frontier model (Gemini 3.1 Pro), while the base model now ranks #14.
Top models on the Mainstream tier
Top models on the Mainstream tier (ACT, DBT, EFT, IFS, Psychodynamic, Polyvagal, AEDP). On the seven widely-represented modalities whose intervention vocabulary lives in many corpora, GPT 5.4 leads (81.5%); Icarus 8.2 (s3160) reaches 76.2%, a +7.9pp lift over its base (68.3%) — the curriculum's largest tier effect. Two 8.2 checkpoints are shown (s2780 74.4%, s3160 76.2%), illustrating continued mainstream gains across training.
Ontology-specific precision leaderboard, zoomed
Ontology-specific lexical and conceptual precision leaderboard (zoomed). Models grouped by deployment tier: Frontier API, Local 229B, and Compact Frontier API. Icarus 8.2 (s3160, 70.8%) ranks #7 overall, +6.1pp above its base model (64.7%, which now ranks #14), lifting it past Grok 4.1 FR and Haiku 4.5; an earlier checkpoint (s2780) edges it at #6 (71.0%).
Precision by modality tier and model
Ontology-specific lexical and conceptual precision by modality tier and model. Three-tier generalization analysis: Mainstream (well-represented modalities), Specialized (less common), and Novel (ontology-specific). The DAPT training effect is positive on every tier and largest on Mainstream (+7.9pp), with equal +5.3pp lifts on Specialized and Novel. On the Novel tier the lift carries 8.2 to #5 in the frontier field — past three models including Sonnet 4.6 (72.3% vs. 69.3%) — the modalities most uniquely aligned with the DAPT curriculum (attachment dynamics, contemplative presence, humanistic/poetic frameworks).
PTAB per-modality precision, base vs Icarus 8.2
PTAB v0.1.3 ontology-specific lexical and conceptual precision: Base MiniMax M2 vs. Icarus 8.2 by modality (output-only). Left: per-modality precision on structured clinical output with think traces stripped, showing separation between base (blue) and 8.2 (red) across most of the ontology set. Right: delta from base (amplified) — red outward = DAPT improved, inward = regressed. DAPT selectively improved modalities aligned with the training curriculum — notably IFS (+18.9pp, its single largest gain), Contemplative I (+10.5pp), EFT (+10.3pp), DBT (+9.4pp), and Humanistic Poetics (+6.0pp, from an already-high ~80% base) — while leaving a few essentially unchanged (Contemplative II, Somatic Trauma at ±0pp) and slightly regressing others (Psychodynamic −3.9pp, AEDP −2.7pp, Polyvagal −2.5pp). This redistribution — gains concentrated on the phenomenological, parts-based, and contemplative ontologies most emphasized in training rather than spread uniformly — is consistent with the hypothesis that LoRA adaptation reorganizes attention toward trained clinical domains rather than injecting uniformly distributed knowledge.
Auto-scored benchmark comparison across 11 aggregates
Auto-scored benchmark comparison across 11 benchmark aggregates, with the two Chinese-language EmoBench subtasks highlighted as hatched pairs: Base MiniMax M2 (blue) vs. Icarus 8.2 DAPT (red). Green/gray delta labels indicate improvement (≥+5%) or stability (< ±5%). Seven of eleven aggregate benchmarks show improvement; none decline beyond −4%. The full EmoBench aggregate improves by +9% across English and Chinese EU/EA items, while the highlighted Chinese subtasks show +4.0% on EU–Chinese and +21.4% on EA–Chinese. This makes visible the largest Chinese-language EmoBench effect alongside EQ-Bench v2 emotion intensity (+50.3%), OpenToM attitude inference (+47%), and ChnSentiCorp Chinese sentiment (+19.2%) — despite the DAPT corpus being entirely English.
Local differentiation at the LoRA boundary (L41), base vs Icarus 8.2
Local differentiation at the LoRA boundary (L41) — base versus Icarus 8.2. Schematic relief of HDBSCAN subclusters from the per-token L41 hidden states of Icarus 8.2's Tier-3e therapist responses, with the same response text passed through both models under an identical clustering threshold. Each mound is one dense token-cluster: its height is how many tokens concentrate there (normalised within each panel), the number of mounds is how finely the model differentiates the region, and how close two mounds sit reflects how similarly the model represents those motifs (metric-MDS of the subcluster centroids' cosine distances in the original 3,072-D L41 space, stress 0.173). Left (base): those activations collapse into a single densely-packed mass (generic punctuation and function words) alongside one 44-token speck — 2 subclusters in total (base's structure here is entangled within one connected mass, not absent — see the cross-representation control in §8.4.2). Right (Icarus 8.2): the same tokens resolve into 76 subclusters whose dominant peaks correspond to recognisable therapeutic content — deprivation & inquiry, integration, aspiration, noticing & attending, therapeutic action, and emotion (the 15 most legible are named). Because height is normalised per panel, the contrast is structural — one mass versus many distinct peaks — not absolute height.
Affect geometry at the commitment layer L61, base vs Icarus 8.2
Affect geometry on identical affect-dense text (L61, the commitment layer). Affect tokens only, sub-clustered and raised as mounds coloured by affect subcategory (layout = metric-MDS of sub-cluster centroids; heights per-panel-normalised). Left (base model): affect concentrates into a narrow, guarded register — two meta-reflective regions (thinking about feeling) and one of felt-safety; clinical and at a remove. Right (Icarus 8.2 model): the same tokens resolve into a wide field of shared felt experience — grief, pain, trauma/shutdown, fear, longing, felt-process. Base reflects on feeling and reassures; Icarus feels with. At the LoRA boundary (L41) the two are comparable on affect; the gap is a downstream, commitment-time property — the affective channel of the specialization-to-commitment arc.
Emotional complexity across zoom at L61, base vs Icarus 8.2
Emotional complexity across zoom at L61 (the commitment layer). One layer downstream from the boundary — and the contrast inverts. Base now collapses to 2–3 affect regions across all resolutions while Icarus 8.2 climbs 4 → 10 → 13. The affective parity present at the boundary (L41) is lost by base at the layer where the model commits to its response; Icarus carries the differentiation through — the affective edge is a downstream, commitment-time property.

Music

compositions · instruments · objects

Chamber Music

String Quartet No. 1

for two violins, viola, and cello

0:00 / 0:00

Hesitance, After Patience - Part 1

for cello trio

0:00 / 0:00

Hesitance, After Patience - Part 2

for cello trio

0:00 / 0:00

Lovers Fall in Love

for harp & two violins

0:00 / 0:00

Cellos & Violins

0:00 / 0:00

Harp & Pitched Percussion (Soundtrack)

0:00 / 0:00

Podcast Opening

0:00 / 0:00

Full Orchestra (Ballet)

(Borrowed grace: the movement of dancers I admire, intentionally carried into my video edit of a ballet I wrote.)

(read along with a rough—and not-yet-engraved—draft of the score here:)

click any page to zoom in

Chess

Multiplayer Chess Boards — Applied Geometry & Fabrication

Tessellating non-standard polygons into playable boards for 3+ players.
all boards are designed for full-sized triple weighted regulation Staunton pieces

Chess board design sketch Chess board design sketch
Three-player chess board Four-player chess board Four-player chess board in red Chess board close-up Chess board close-up Chess board close-up Chess board close-up Chess board close-up Chess board close-up Chess board close-up Chess board close-up Chess board close-up

Inspiration

"Social scientists must know the technical specifics of the AI system we are studying."— Giada Pistilli
What lies behind AGI: ethical concerns related to LLMs
"Everybody's going for more intelligence….We should be very careful to make…beings that care about us….We can still do that….What we want to create is a new kind of being that cares primarily about people…And we're not gonna get that by invisible hand…What we want to do is be designing them so that they'll be nice to us…And I think we should put a lot of effort into that…"— Geoffrey Hinton
"We must build AI for people; not to be a person… I want to create AI that makes us more human, that deepens our trust and understanding of one another, and that strengthens our connections to the real world."— Mustafa Suleyman, CEO of Microsoft AI
"AI systems are never going to truly support human thinking and learning if they're not able to participate in these multimodal exchanges using all of the information channels that humans naturally use."— Judy Fan
"We must shape the values of AI with humanity's common values, and make good use of AI technologies to increase understanding, tolerance, exchanges, and sharing among all civilizations. We should tend to the garden of civilizations with great care, to ensure that beauty of each civilization is appreciated and shared."— Xi Jinping
"Whatever we may, you know, feel like, being a machine learner is all about, you know, you think about the data. You look at it. You come up with hypotheses. Maybe equivariance is important. You try it. You measure. 9 times out of 10, you find out you're wrong. Right? If you're wrong 9 times out of 10, you're a very successful machine learner. You're incredibly productive. And you build this local intuition. You build this notion of what the problem needs and how it works, and you build, in my view, a kind of science local to your area, kind of a local manifold of ideas."— John Jumper
"Many people have experienced extraordinary moments of revelation, creative inspiration, compassion, fulfillment, transcendence, love, beauty, or meditative peace… the 'space of what is possible to experience' is very broad and a larger fraction of people's lives could consist of these extraordinary moments."— Dario Amodei, CEO of Anthropic
Machines of Loving Grace
"The things that make us human will become much more important instead of much less important…"— Daniela Amodei, Co-founder & President of Anthropic
"We are just getting started….[And this work] still needs soul and craft, the human taste in whatever you are building."— Demis Hassabis, CEO of DeepMind
"...be more of an artist, look at things from a different direction, be able to build something unique..."— Alex Karp, CEO of Palantir
"I think their rate of discovery could be increased by 10x or more if there were a lot more talented, creative researchers… the returns to intelligence are high for these discoveries, and everything else in biology and medicine mostly follows from them."— Dario Amodei, CEO of Anthropic
Machines of Loving Grace

AI is exposing new human bottlenecks

When one part of a workflow becomes dramatically faster, the constraint simply moves elsewhere. Organizations then need more people to remove the next bottleneck and capture the value AI has unlocked.

Automation does not eliminate the need for human contribution. It often reveals how much more could be accomplished with it.

AI fluency will become one of the world’s most valuable skills

The emerging divide will not be between humans and agents. It will be between people who are highly fluent with AI and those who are not.

AI-fluent people will not be 10% more productive. In some forms of work, they could be 50x or even 100x more effective. They will imagine better uses for AI, orchestrate agents and apply judgment where machines still fall short.

These people will be scarce and enormously valuable.

— Jeetu Patel

“The task of converting observations into numbers is the hardest of all, the last task rather than the first thing to be done, and it can be done only when you have learned, beforehand, a great deal about the observations themselves. You can, to be sure, achieve a very deep understanding of nature by quantitative measurement, but you must know what you are talking about before you can begin applying the numbers for making prediction…

“The risks of untoward social consequences in work of this kind are considerable. It is as important—and as hard—to learn when to use mathematics as how to use it, and this matter should remain high on the agenda of consideration for education in the social and behavioral sciences…

“The instruments devised for approaching nature… became increasingly precise and powerful, carrying… triumph after triumph, and setting the stage for the great revolution… There is no doubt about it: measurement works when the instruments work, and when you have a fairly clear idea of what it is that is being measured, and when you know what to do with the numbers when they tumble out…

“We have a wilderness of mystery to make our way through in the centuries ahead, and we will need science for this but not science alone. Science will, in its own time, produce the data and some of the meaning in the data, but never the full meaning. For getting a full grasp, for perceiving real significance when significance is at hand, we shall need minds at work from all sorts of brains outside the fields of science, most of all the brains of poets, of course, but also those of artists, musicians, philosophers, historians, writers in general… [There will be those, too, who] for their part, might take considerable satisfaction watching their scientific colleagues confess openly to not knowing everything about everything. And the poets, on whose shoulder the future rests, might, late nights, thinking things over, begin to see some meanings that elude the rest of us. It is worth a try.”

— Lewis Thomas, 1980
Download
Loading map…