Author Archives: Victoria Krakovna

Moving on from community living

After 7 years at Deep End (and 4 more years in other group houses before that), Janos and I have moved out to live near a school we like and some lovely parks. The life change is bittersweet – we will miss living with our friends, but also look forward to a logistically simpler life with our kids. Looking back, here are some thoughts on what worked and didn’t work well about living in a group house with kids.

Pros. There were many things that we enjoyed about living at Deep End, and for a long time I couldn’t imagine ever wanting to leave. We had a low-effort social life – it was great to have spontaneous conversations with friends without arranging to meet up. This was especially convenient for us as new parents, when it was harder to make plans and get out of the house, particularly when we were on parental leave. The house community also made a huge difference to our wellbeing during the pandemic, because we had a household bubble that wasn’t just us.

We did lots of fun things together with our housemates – impromptu activities like yoga / meditation / dancing / watching movies, as well as a regular check-in to keep up on each other’s lives. We were generally more easily exposed to new things – meeting friends of friends, trying new foods or activities that someone in the house liked, etc. Our friends often enjoyed playing with the kids, and it was helpful to have someone entertain them while we left the living room for a few minutes. Our 3 year old seems more social than most kids of the pandemic generation, which is partly temperament and partly growing up in a group house.

Cons. The main issue was that the group house location was obviously not chosen with school catchment areas or kid-friendly neighbourhoods in mind. The other downsides of living there with kids were insufficient space, lifestyle differences, and extra logistics (all of which increased when we had a second kid).

Our family was taking up more and more of the common space – the living room doubled as a play room and a nursery, so it was a bit cramped. With 4 of us (plus visiting grandparents) and 4 other housemates in the house, the capacity of the house was maxed out (particularly the fridge, which became a realm of mystery and chaos). I am generally sensitive to clutter, and having the house full of our stuff and other people’s stuff was a bit much, while only dealing with our own things and mess is more manageable.

Another factor was a mismatch in lifestyles and timings with our housemates, who tended to have later schedules. They often got home and started socializing or heading out to evening events when we already finished dinner and it was time to put the kids to bed, which was FOMO-inducing at times. Daniel enjoyed evening gatherings like the house check-in, but often became overstimulated and was difficult to put to bed afterwards. The time when we went to sleep in the evening was also a time when people wanted to watch movies on the projector, and it made me sad to keep asking them not to.

There were also more logistics involved with running a group house, like managing shared expenses and objects, coordinating chores and housemate turnover. Even with regular decluttering, there was a lot of stuff at the house that didn’t belong to anyone in particular (e.g. before leaving I cleared the shoe rack of 9 pairs of shoes that turned out to be abandoned by previous occupants of the house). With two kids, we have more of our own logistics to deal with, so reducing other logistics was helpful.

Final thoughts. We are thankful to our housemates, current and former, for all the great times we had over the years and the wonderful community we built together. Visiting the house after moving out, it was nice to see the living room decked out with pretty decorations and potted plants and not overflowing with kid stuff – it reminded me of what the house was like when we first started it. Without the constraints of children living at the house, I hope to see Deep End return to its former self as a social place with more events and gatherings, and we will certainly be back to visit often.

It is a big change to live on our own after all these years. We moved near a few other friends with kids, which will be fun too. We are enjoying our own space right now, though we are not set on living by ourselves indefinitely. We might want to live with others again in the future, but probably with 1-2 close friends rather than in a big group house.

2023-24 New Year review

2023 review

Life updates

We received a special gift for New Year’s – Michael (“Misha”) arrived just in time to be born in 2023! Daniel is already getting the hang of rocking his brother and singing him lullabies.

Continue reading →

Retrospective on my posts on AI threat models

When discussing AI risks, talk about capabilities, not intelligence

2 Replies

Public discussions about catastrophic risks from general AI systems are often derailed by using the word “intelligence”. People often have different definitions of intelligence, or associate it with concepts like consciousness that are not relevant to AI risks, or dismiss the risks because intelligence is not well-defined. I would advocate for using the term “capabilities” or “competence” instead of “intelligence” when discussing catastrophic risks from AI, because this is what the concerns are really about. For example, instead of “superintelligence” we can refer to “super-competence” or “superhuman capabilities”.

When we talk about general AI systems posing catastrophic risks, the concern is about losing control of highly capable AI systems. Definitions of general AI that are commonly used by people working to address these risks are about general capabilities of the AI systems:

PASTA definition: “AI systems that can essentially automate all of the human activities needed to speed up scientific and technological advancement”.
Legg-Hutter definition: “An agent’s ability to achieve goals in a wide range of environments”.

We expect that AI systems that satisfy these definitions would have general capabilities including long-term planning, modeling the world, scientific research, manipulation, deception, etc. While these capabilities can be attained separately, we expect that their development is correlated, e.g. all of them likely increase with scale.

There are various issues with the word “intelligence” that make it less suitable than “capabilities” for discussing risks from general AI systems:

Anthropomorphism: people often specifically associate “intelligence” with being human, being conscious, being alive, or having human-like emotions (none of which are relevant to or a prerequisite for risks posed by general AI systems).
Associations with harmful beliefs and ideologies.
Moving goalposts: impressive achievements in AI are often dismissed as not indicating “true intelligence” or “real understanding” (e.g. see the “stochastic parrots” argument). Catastrophic risk concerns are based on what the AI system can do, not whether it has “real understanding” of language or the world.
Stronger associations with less risky capabilities: people are more likely to associate “intelligence” with being really good at math than being really good at politics, while the latter may be more representative of capabilities that make general AI systems pose a risk (e.g. manipulation and deception capabilities that could enable the system to overpower humans).
High level of abstraction: “intelligence” can take on the quality of a mythical ideal that can’t be met by an actual AI system, while “competence” is more conducive to being specific about the capability level in question.

It’s worth noting that I am not suggesting to always avoid the term “intelligence” when discussing advanced AI systems. Those who are trying to build advanced AI systems often want to capture different aspects of intelligence or endow the system with real understanding of the world, and it’s useful to investigate and discuss to what extent an AI system has (or could have) these properties. I am specifically advocating to avoid the term “intelligence” when discussing catastrophic risks, because AI systems can pose these risks without possessing real understanding or some particular aspects of intelligence.

The basic argument for catastrophic risk from general AI has two parts: 1) the world is on track to develop generally capable AI systems in the next few decades, and 2) generally capable AI systems are likely to outcompete or overpower humans. Both of these arguments are easier to discuss and operationalize by referring to capabilities rather than intelligence:

For #1, we can see a trend of increasingly general capabilities, e.g. from GPT-2 to GPT-4. Scaling laws for model performance as compute, data and model size increase suggest that this trend is likely to continue. Whether this trend reflects an increase in “intelligence” is an interesting question to investigate, but in the context of discussing risks, it can be a distraction from considering the implications of rapidly increasing capabilities of foundation models.
For #2, we can expect that more generally capable entities are likely to dominate over less generally capable ones. There are various historical examples of this, e.g. humans causing other species to go extinct. While there are various ways in which other animals may be more “intelligent” than humans, the deciding factor was that humans had more general capabilities like language and developing technology, which allowed them to control and shape the environment. The best threat models for catastrophic AI risk focus on how the general capabilities of advanced AI systems could allow them to overpower humans.

As the capabilities of AI systems continue to advance, it’s important to be able to clearly consider their implications and possible risks. “Intelligence” is an ambiguous term with unhelpful connotations that often seems to derail these discussions. Next time you find yourself in a conversation about risks from general AI where people are talking past each other, consider replacing the word “intelligent” with “capable” – in my experience, this can make the discussion more clear, specific and productive.

(Thanks to Janos Kramar for helpful feedback on this post.)

Near-term motivation for AI alignment

2022-23 New Year review

2 Replies

This is an annual post reviewing the last year and setting goals for next year. Overall, this was a reasonably good year with some challenges (the invasion of Ukraine and being sick a lot). Some highlights in this review are improving digital habits, reviewing sleep data from the Oura ring since 2019 and calibration of predictions since 2014, an updated set of Lights habits, the unreasonable effectiveness of nasal spray against colds, and of course baby pictures.

2022 review

Life updates

I am very grateful that my immediate family is in the West, and my relatives both in Ukraine and Russia managed to stay safe and avoid being drawn into the war on either side. In retrospect, it was probably good that my dad died in late 2021 and not a few months later when Kyiv was under attack, so we didn’t have to figure out how to get a bedridden cancer patient out of a war zone. It was quite surreal that the city that I had visited just a few months back was now under fire, and the people I had met there were now in danger. The whole thing was pretty disorienting and made it hard to focus on work for a while. I eventually mostly stopped checking the news and got back to normal life with some background guilt about not keeping up with what’s going on in the homeland.

Work

My work focused on threat models and inner alignment this year:

Made an overview talk on Paradigms of AI alignment: components and enablers and gave the talk in a few places.
Coauthored Goal Misgeneralization: why correct rewards aren’t enough for correct goals paper and the associated DeepMind blog post
Did a survey of DeepMind alignment team opinions on AGI ruin arguments, which received a lot of interest on the alignment forum.
Wrote a post on Refining the Sharp Left Turn threat model
Contributed to DeepMind alignment posts on Clarifying AI x-risk and Threat model literature review
Coauthored a prize-winning submission to the Eliciting Latent Knowledge contest: Route understanding through the human ontology.

Continue reading →

Refining the Sharp Left Turn threat model

2 Replies

(Coauthored with others on the alignment team and cross-posted from the alignment forum: part 1, part 2)

A sharp left turn (SLT) is a possible rapid increase in AI system capabilities (such as planning and world modeling) that could result in alignment methods no longer working. This post aims to make the sharp left turn scenario more concrete. We will discuss our understanding of the claims made in this threat model, propose some mechanisms for how a sharp left turn could occur, how alignment techniques could manage a sharp left turn or fail to do so.

Claims of the threat model

What are the main claims of the “sharp left turn” threat model?

Claim 1. Capabilities will generalize far (i.e., to many domains)

There is an AI system that:

Performs well: it can accomplish impressive feats, or achieve high scores on valuable metrics.
Generalizes, i.e., performs well in new domains, which were not optimized for during training, with no domain-specific tuning.

Generalization is a key component of this threat model because we’re not going to directly train an AI system for the task of disempowering humanity, so for the system to be good at this task, the capabilities it develops during training need to be more broadly applicable.

Some optional sub-claims can be made that increase the risk level of the threat model:

Claim 1a [Optional]: Capabilities (in different “domains”) will all generalize at the same time

Claim 1b [Optional]: Capabilities will generalize far in a discrete phase transition (rather than continuously)

Claim 2. Alignment techniques that worked previously will fail during this transition

Qualitatively different alignment techniques are needed. The ways the techniques work apply to earlier versions of the AI technology, but not to the new version because the new version gets its capability through something new, or jumps to a qualitatively higher capability level (even if through “scaling” the same mechanisms).

Claim 3: Humans can’t intervene to prevent or align this transition

Path 1: humans don’t notice because it’s too fast (or they aren’t paying attention)
Path 2: humans notice but are unable to make alignment progress in time
Some combination of these paths, as long as the end result is insufficiently correct alignment

Continue reading →

Paradigms of AI alignment: components and enablers

1 Reply

(This post is based on an overview talk I gave at UCL EA and Oxford AI society (recording here). Cross-posted to the Alignment Forum. Thanks to Janos Kramar for detailed feedback on this post and to Rohin Shah for feedback on the talk.)

This is my high-level view of the AI alignment research landscape and the ingredients needed for aligning advanced AI. I would divide alignment research into work on alignment components, focusing on different elements of an aligned system, and alignment enablers, which are research directions that make it easier to get the alignment components right.

Alignment components
- Outer alignment
- Inner alignment
Alignment enablers

You can read in more detail about work going on in these areas in my list of AI safety resources.

Continue reading →

2021-22 New Year review

1 Reply

This was a rough year that sometimes felt like a trial by fire – sick relatives, caring for a baby, and the pandemic making these things more difficult to deal with. My father was diagnosed with cancer and passed away later in the year, and my sister had a sudden serious health issue but is thankfully recovering. One theme for the year was that work is a break from parenting, parenting is a break from work, and both of those things are a break from loved ones being unwell.

I found it hard to cope with all the uncertainty and stress, and this was probably my worst year in terms of mental health. There were some bright spots as well – watching my son learn many new skills, and lots of time with family and in nature. Overall, I look forward to a better year ahead purely based on regression to the mean.

2021 review

Life updates

My father, Anatolij Krakovny, was diagnosed with late-stage lung cancer in January with a grim prognosis of a few months to a year of life. This came out of nowhere because he’s always been healthy and didn’t have any obvious risk factors. We researched alternative treatments to the standard chemotherapy and arranged additional tests for him but didn’t find anything promising.

We went to Ukraine to visit him in February and he was happy to meet his grandson. We were worried about the covid risks of traveling with little Daniel but concluded that they were low enough, and thankfully we were allowed to leave the UK though international travel was not generally permitted.

My dad seemed to have a remission in the summer, and we considered visiting him in June, but he told us not to come because of the covid situation in Ukraine. Unfortunately we listened to him and didn’t go (this would have been a good opportunity to spend time with him while he was still doing well).

We spent most of the summer in Canada, with grandparents taking care of Daniel. This was a relaxing time with family and nature, until my sister had a sudden life-threatening health problem and was in and out of hospital with a lot of uncertainty around recovery. This also came out of the blue with no obvious risk factors present. She is feeling better now and doctors expect a full recovery, which we are very grateful for.

In November, my dad had a sudden relapse, and we went to Ukraine again. Once there we realized that the public health system wasn’t taking good care of him (they were mostly swamped with covid) and we had to find a private hospital to take him in. He was already in pretty bad shape and died two weeks later, but I’m glad we managed to see him and help him in some way.

Continue reading →

Reflections on the first year of parenting

1 Reply

The first year after having a baby went by really fast – happy birthday Daniel! This post is a reflection on our experience and what we learned in the first year.

Grandparents. We were very fortunate to get a lot of help from Daniel’s grandparents. My mom stayed with us when he was 1 week – 3 months old, and Janos’s dad was around when he was 4-6 months old (they made it to the UK from Canada despite the pandemic). We also spent the summer in Canada with the grandparents taking care of the baby while we worked remotely.

We learned a lot about baby care from them, including nursery rhymes in our respective languages and a cool trick for dealing with the baby spitting up on himself without changing his outfit (you can put a dry cloth under the wet part of the outfit). I think our first year as parents would have been much harder without them.

Continue reading →