Showing posts with label road testing. Show all posts
Showing posts with label road testing. Show all posts

Sunday, July 16, 2023

AV Safety and the False Dilemma Fallacy

The current AV company messaging strategy is a classic case of a false dilemma fallacy. They frame the situation as a choice between continued human drivers killing people (without statistical context) vs. immature robotaxis who don't drink and drive (but make other mistakes). (Wikipedia: False dilemma)


The recent Cruise ad in particular is a plainly ridiculous doubling-down on the industry's long discredited propaganda playbook.

Cruise NY Times ad: https://twitter.com/kvogt/status/1679517290847694848


Analysis of AV industry playbook: https://www.eetimes.com/autonomous-vehicle-myths-the-dirty-dozen/


A more reasonable message would be cities need robotaxis for <reasons> and robotaxi companies will use <defined, balanced metrics, stated in advance rather than cherry picked later> to show they are no worse than human drivers during development, with monthly report card disclosures. Improved safety comes later -- we all hope.


Here is where things really stand:

  • It is too early to know whether current robotaxi technology is safer than human drivers for fatalities. The industry is stringing us along hoping they can show they are safe over time (starting now, but not really there yet).
  • Non-autonomous technology (AEB) is making far more of a contribution right now -- but is missing from the false dilemma.
  • Public transit (much safer) also not in the discussion.
  • Improving road safety (speed limits, traffic calming, etc.) also not in the discussion. Also missing are specific pedestrian and cyclist safety improvements. While we're at it, seat belts, drunk driving, and motorcycle safety measures.
  • The messaging from both sides (parts of the SF govt and especially Cruise) on crashes does not address factors required for a reasonable comparison. (ODD, baseline driver population, etc.)
  • It is clear that vehicles from both Waymo and Cruise are creating public road disruption. There is no excuse for impeding emergency responders just to get "Look Ma, No Driver!!" optics.
  • The technology can be advanced by continuing to test while having human drivers or in-car valets (employee in the front passenger seat) to mitigate problems. Their Safety Management System should include a step of adding/removing in-car vehicle supervisors until issues that cause public disruption are shown to be resolved in deployment.
  • Cruise in particular needs to get more diligent about pre-deployment testing. There is simply no excuse for rear-ending a Muni bus due to a software defect in an uncrewed vehicle that occurred in a pretty normal situation. Waymo isn't perfect, but their failures are more at the edge.
  • The public outrage is entirely self-inflicted by companies due to their exploitation of the municipal preemption clause in state-level regulations rather than being responsible road users. Playing the "we should forgive their drivers who are still learning" card has worn out its welcome.


Some might want to point out that some companies are worse actors than others, but all the companies have their issues. (For example, good work by the Waymo safety team is hurt by their government relations breathless safety hype messaging.)


And the reality is that a crash or adverse news for one company hurts them all.


Wednesday, December 7, 2022

SCSC Talk: Bootstrapping Safety Assurance

Bootstrapping Safety Assurance

Abstract:
The expense and general impracticability of doing enough real-world testing to demonstrate safety for autonomous systems motivates finding some sort of shortcut. A bootstrapped testing approach is often proposed, using evidence from initial mishap-free testing to argue that continued testing is safe enough. In this talk I'll explain why pure bootstrapping based on testing exposure as well as arguments involving "probably perfect" bootstrapping expose public road users to undue risk. Moreover, phased deployments often used to argue safe update release have the same problem. An approach that bootstraps on the safety case rather than on vehicle testing is proposed as a potentially better alternative. While the examples given involve autonomous ground vehicles, the principles involved apply to any argument that safety will be demonstrated via a bootstrap testing process.

This talk was recorded as part of the SCSC Future of Testing for Safety-Critical Systems seminar on Dec. 1, 2022.
Talks and videos are available here (access with paid annual club membership):  https://scsc.uk/e966prog

Free public-access copy of slides here: 




Friday, November 4, 2022

What Happens After The Industry's JohnnyCab Adventure?

Recent news has people questioning whether autonomous vehicles are viable. Promises that victory is right around the corner are too optimistic. But it's far too early to declare defeat. There is a lot to process. But a pressing technology roadmap question is: if robotaxis aren't really the answer, what happens next for passenger vehicles?

Ahhhnold starts an ill-fated JohnnyCab trip in Total Recall (older version)

robotoxi ride that did not go as expected. (Total Recall 1990)

We can expect OEMs to double down on their Level 2è2+è3 strategies. But there is a very real risk of a race to the bottom as companies scramble to deploy shiny automation technology while skipping over reasonable, industry-created safety practices. A subtle, yet crucial, point is that asking a driver to supervise a driver assistance system is a completely different situation than putting a civilian car owner in the position of being an untrained autonomous vehicle test driver.

Where are we now?

The recent demise of Argo AI has made it crystal clear that Big Auto is pivoting away from robotaxis. Big Auto has kept plugging away at driver assistance systems while hedging its bets by taking stakes in Level 4 robotaxi companies. Those Level 4 companies were carefully kept at arms length against the day it was time to close out those hedges. And here we are.

Other companies continue to work on robotaxis -- most notably Waymo, Cruise, and some players in China that are carrying passengers on public roads. Waymo seems unconstrained by runway length. Presumably Cruise has a year-end milestone as they did in 2020 and 2021, and we'll see how that goes. Meanwhile Tesla continues to celebrate anniversaries of its promise to put a million robotaxis on the road within a year. Heavy trucks, parcel delivery, and low speed shuttles are also part of the mix, but face their own challenges that are beyond the scope of this discussion.

For now let's entertain the possibility that robotaxis are not going to happen anytime soon. It seems likely to take years of hammering away at the heavy tail of edge cases to get there. Sure, some of us can get a demo ride in a Taxi of Tomorrow in a couple cities. But let's assume that robotaxis are more of a Disney-esque experience than a practical tool for profitable mobility at scale anytime soon.

What happens next?

We can expect to see OEMs double down on evolutionary strategies. Ford has pretty much said as much, promising a Level 3 system is on the way. It will probably be talked about as climbing up the SAE Levels from 2 to 3 to 4, which has been their narrative all along. That narrative has significant issues because it makes the common mistake of using SAE levels as a progressive ladder to climb -- which it is definitely not. But it will be the messaging framework nonetheless.

In the mix are several options:

  • Level 2 highway cruising systems that are already on the road
  • Level 2+ add-ons, but with the driver still responsible for safety
  • Traffic jam pilot as the first step to drivers taking attention off the road
  • Level 3/4 capabilities beyond traffic jams (harder than it might seem)
  • Abuse of Level 2/2+ designations to evade regulation

We'll see all of these in play. Each option comes with its own challenges.

Level 2 systems: highway cruising

The general idea is that first you build an SAE Level 2 system that has a speed+steering cruise control capability (automated speed control and automated lane keeping). This is intended for use on highways and other situations involving well behaved, in-lane travel. This functionality is already widely available on new high-end vehicles. The driver is responsible for continuously monitoring the car's performance and the road conditions, and intervening instantly if necessary -- whether the car asks for intervention or not.

A significant challenge for these systems is ensuring that drivers remain attentive in spite of inevitable automation complacency. This requires a highly capable driver monitoring system (DMS) that, for example, tracks where the driver is looking to ensure constant monitoring of the situation on the roadway.

Performance standards for monitoring driver attentiveness are in their infancy, and capabilities of the DMS on different vehicles varies dramatically. It is pretty clear that a camera-based system is much more effective than steering wheel touch sensors. And a camera that can see the driver's eye gaze direction, including having infrared emitters for night operation, is likely to be more effective than one that can't.

Another crucial safety feature that should be implemented is restricting operation of Level 2 features to the conditions they are designed to be safe in. The SAE J3016 taxonomy standard makes this totally optional -- which is one of the many reasons J3016 is not a safety standard, and why no government should write the SAE Levels into their safety regulations. But if the automation is supposed to only be used on divided highways with no cross-streets, then it should refuse to engage unless that is the road type it is being used on. Some companies restrict their vehicle to operation on specific roads, but others do not.

NTSB has made a number of recommendations to improve DMS capability and related Level 2 functionality, but the industry is still working on getting there.

Vehicle automation proponents have been beating the drum of potential safety benefit for years -- to the point that many take safety benefits as an article of faith. However, the reality is that there is no credible evidence that Level 2 capabilities make systems safer. There is plenty of propaganda to be sure, but much of that is based on questionable comparisons of very capable high-end vehicles with AEB and five-star crash test results vs. fleet-average 12 year old cars without all that fancy safety technology. Comparisons also tend to involve apples-meets-oranges different operational scenarios. The results are essentially meaningless, and simply provide grist for the hype machine.

At best, from available information, Level 2 systems are a safety-neutral convenience feature. At worst, AEB is compensating for a moderate overall safety loss from Level 2 features, contributing to an ever-increasing number of fatalities that might be associated with disturbing patterns such as crashing into emergency responders. However, nobody really knows the extent of any problems because of extreme resistance to sharing data by car companies that would reveal the true safety situation. (NHTSA has mandated crash reporting for some crashes, but there is no denominator of miles driven, so crash rates cannot be determined in a reliable way using publicly available data.) 

If the safety benefits were really there, one imagines that these companies full of PhDs who know how to write research papers would be falling all over themselves to publish rigorously stated safety analyses. But all we're seeing is marketing puffery, so we should assume the safety benefit is not yet there (or at least potential safety benefits are not yet supported by the actual data).

In the face of these questions about automation safety and pervasive industry safety opacity, the industry's plan is to ... add even more automation.

Level 2+ systems

"Level 2+" is a totally made up marketing term. The SAE J3016 standard says this term is invalid. But it gets used anyway so let's go with it and try to see what it might mean in practice

Once you have a vehicle that stays in its lane and avoids most other vehicles, companies feel competitive pressure to add more capability. Features such as lane changing, automatically merging into traffic, and automatically detecting traffic lights might help drivers with workload. And sure, that sounds nice if it can be done safely.

The question is what happens to driver attention as automation features get added and automation performance improves. There is plenty of reason to believe that as the driver perceives automation quality is improving, they will struggle to remain engaged with vehicle operation. They will succumb to driver complacency. The better the automation, the worse the problem, as illustrated by this conceptual graph:

As autonomous features become more reliable net safety can decrease due to driver complacency

A significant complication is that the driver has to monitor not only the road conditions, but also the car's behavior to determine if the car will do the right thing or not. This can be pretty tricky, especially if it is difficult to know if the car "sees" something happening on the road or not. What if there is a firetruck stopped on the highway in front of you? Will the car detect and stop, or run right into it? How long do you wait to find out? What if you wait just a little too long and suffer a crash? Monitoring automation behavior can be tricky, and is a much different task than simply driving. Assuming that a competent driver is also a competent monitor is a bit of a leap of faith. As you add more complex automation behavior, the driver's ability to track what is and is not supposed to be handled and compare that to what the vehicle is doing can easily be overwhelming. 

Monitoring cross-traffic can be especially difficult. How can the supervising driver know their car is about to accelerate out across traffic to make a left turn when there is an oncoming car?  Pressing the brake quickly after the car has already started lunging into traffic as a safety supervising driver is an entirely different safety proposition than waiting for a driver to command speed when the road is clear. The human/computer interface implications here are tricky.

Car companies will continue to pile on more automation features. At least some of them will continue to improve DMS capability. There are many human/machine interface issues here to resolve, and the outcome will in large part depend on non-linear effects related to often the human driver needs to intervene quickly and correctly to ensure safety. How that will work out in real world conditions remains to be seen. 

The moral hazard of blaming the driver

A significant issue with Level 2 and 2+ systems is that the driver is responsible for safety, but might not be put into a position to succeed. There are natural limits to human performance when supervising automation, and we know that it is difficult for humans to do well at this task. We should not ask human drivers to display superhuman capabilities, then blame them for being mere mortals. Driver monitoring might help, but we should respect the limits to how much it can help.

There is a temptation to blame the driver for any crash involving automation technology. However, doing so is counterproductive. A crash is a crash. Blame does not change the fact that a crash happened. Blame is a fairly ineffective deterrent to other drivers slacking off in general, and is wholly ineffective at converting normal humans into super-humans. If using a Level 2 or 2+ system results in a net increase in crashes, blaming the drivers won't bring the increased fatalities back to life -- even if we bring criminal charges.

If real-world crashes increase with the use of driver automation (compared to a comparably equipped vehicle in the same conditions without driver automation), they should be considered  unreasonably dangerous. Something would need to be done about such systems to change the system, human behavior, or both. Changing human behavior via "education" is usually what is attempted, and almost never works. The human/computer interface and the feature set are much more likely what need to change.

As a simple example of how this might play out, let's say a car company runs TV advertisements showing drivers singing and clapping hands while engaging a Level 2/2+ system. (Dialing back the #autonowashing, later ads show only the passengers doing this.) The company should have data showing that even if a worst case driving automation error for their system takes place mid-clap, the driver will be attentive enough to the situation (despite being caught up in the song) and have a quick enough reaction time to get their hands back onto the steering wheel and intervene for safety. Blaming the driver for clapping hands after airing a TV commercial amounts should not be a permissible tactic. 

Note that we are not saying hands-off is inherently unsafe, but rather that permitting (even encouraging) hands-off operation significantly increases the challenge of ensuring practical safety. The data needs to be there to justify safety for whatever operational concept is being deployed and encouraged.

ALKS: a baby Step Toward Level 3

The way the term "Level 3" is being used by almost everyone seldom matches what the SAE J3016 standard actually says. (See Myths #6, #7, #8. #14, #15, etc in this J3016 analysis.) But this is not the place to hash out that incredibly messy situation other than to note that SAE J3016 terminology is not how we should be describing these safety critical driving features at all.

For our purposes, let's assume when a car maker says "Level 3" they mean that the driver can take their eyes off the road, at least for a little bit, so long as they are ready to jump back into the role of driving if the car sounds an alarm for them to do so. 

The industry's baby step toward this vision of Level 3 is Automated Lane Keeping Systems (ALKS) as described by the European standard UNECE Reg. 157. (To its credit, this standard does not mention SAE Levels at all.) The short version is that drivers can take their eyes off the road and let the car do the driving in slow speed situations. In general, this is envisioned as a traffic jam assistant (slow speed, stop-and-go, traffic jam situations). We can expect this to be the first Level 3 step in an envisioned transition from Level 2 to 2+ to 3 strategy that is already said to be deployed at small scale in Europe and Japan.

ALKS might work out well.  Traffic jams on freeways are pretty straightforward compared to a lot of other operational scenarios. Superficially there are a lot of slow moving cars and not much else. Going just a bit deeper, if the jam is due to a crash there will be road debris, emergency responders, and potentially dazed victims walking through traffic. So saying "no pedestrians" or even "pedestrians will be well behaved" is unrealistic near a crash scene. But one can envision this can be managed, so it seems a reasonable first step so long as safety is considered carefully. 

Broader Level 3 safety requirements

Going beyond the strict workings of SAE J3016 (which, if followed to the letter and not exceeded, will almost certainly result in an unsafe vehicle), Level 3 driving safety only works if:

  • The Automated Driving System (ADS) is 100% responsible for safety the entire time it is engaged. Period. Any weasel wording "but the driver should..." will lead to Moral Crumple Zone designs (blaming people for automation failing to operate as advertised). Put another way, if the driver is ever blamed for a crash when a Level 3 automation feature is engaged, it wasn't really Level 3.
  • The ADS needs to determine with essentially 100% reliability when it is time for the driver to intervene. This is one aspect of the most difficult problem in autonomous vehicle design: having the automation know when it doesn't know what to do. The magnitude of this challenge should not be underestimated. Safe Level 3 is not just a little harder than Level 2. The difference is night and day in any but the most benign operational conditions. Machine learning is terrible at handling unknowns (things it hasn't trained on), but recognizing something is an unknown is required to make sure the driver intervenes when the ADS cannot handle the situation.
  • The ADS ensures reasonable safety until the driver has time to respond, despite the fact that something has gone wrong (or is about to go wrong) to prompt the takeover request. This means the ADS needs to keep the car safe at least for a while. If the driver takes a long time to respond, the ADS needs to do something reasonable. In some cases perhaps an in-lane stop is good enough; in others not. (In practice this pushes the ADS arguably to be a very low-end Level 4 system, but we're back to J3016 standards gritty details so let's not even go there. The point is that the driver might take a long time to respond, and the ADS can't simply dump blame on the driver if a crash happens when the going gets tough.)
For ALKS, the main safety plan is to go slowly enough that the car can stop before it is likely to hit anything. That "anything" is predominantly other cars, and sometimes people at an emergency scene. High speed animal encounters, cross traffic, and other situations are ruled out by the narrowly limited operational design domain (ODD), which should be enforced by prohibiting activation in inappropriate situations. The ALKS standard presumes drivers will respond to takeover notifications in 10 seconds, but the vehicle has to do something reasonable even if that does not happen.

More generally, it is OK to take credit for most drivers being able to respond relatively quickly in most situations. It is likely that a more advanced Level 3 system will progressively degrade its operation over a period of many seconds, such as first slowing down, then coming to a stop in the safest place it can reach given the situation.  If a driver falls asleep from boredom, it might take a while for honking horns from other drivers wake them up. Or it might take longer than 10 seconds to regain situational awareness if overly engrossed in a movie they were told was OK to watch while in this operational mode. Human subject studies can be used to claim credit for most drivers intervening relatively quickly  (to the degree that is true) even if all drivers cannot respond that quickly. Credit could be taken for being able to pull over to the side of the road in most situations to reduce the risk of being hit by other vehicles, even if that is not possible in all situations. Safety for human driver takeover is not the result of a single ten second threshold, but rather a stack-up of various levels of degraded operation and probabilities of mishaps in a variety of real-world scenarios.

One might be able to argue net acceptable safety as long as the worst cases are infrequent. If the worst cases happen too often -- including inevitable misuse and abuse -- some sort of redesign will be needed. Again, blaming drivers or "educating" them won't fix a fundamentally unsafe operational concept. Safety must be built into the system, not dodged by blaming drivers for having an unacceptably high crash rate.

Advanced features and test platforms

Once ALKS is in place, the inevitable story is going to be to slowly expand the ODD. For example, if things seem safe, increase the speed at which ALKS can operate. Then let it do on/off ramps. And so on.

The same thinking will be in place for Level 2+ systems. If in-lane cruise control works, add lane changing. Add traffic merging. Let it operate in suburbs. Add unprotected left turns. Next thing you know, we'll be back to saying robotaxis will be here next year -- this time for sure!

While this seductive story sounds promising, it's not going to be easy. Automation complacency combined with limits of drivers to supervise potentially unpredictable autonomy functions will be a first issue to overcome. The second will be confusing a Level 2 driver with an autonomous vehicle testbed driver.

We've already discussed automation complacency. However, it is going to get dramatically worse with more advanced functionality. Humans learn to trust automation far too quickly. If a car correctly makes an unprotected left turn 100 times in a row, any safety driver might well be mentally unprepared to intervene in time when it pulls out in front of oncoming traffic on the 101st turn. Even something as simple as a automation being able to see moving cars but not detecting an overturned truck on the road has led to crashes, presumably due to automation complacency.

While car companies can argue that trained, professional safety drivers might keep testing acceptably safe (e.g., by conforming to the SAE J3018 safety standard -- which mostly is not being done), that does not apply to retail customers who are driving vehicles.

Tesla has led the industry in some good ways by promoting the adoption of electric vehicles. But it has set a terrible precedent by also adopting a policy of having untrained civilian vehicle owners operate as testers for immature automation technology, leading to reckless driving on public roads.

There is already significant market pressure for other companies to follow the Tesla playbook by pushing out immature automation features to their vehicles and telling drivers that they are responsible for ensuring safety. After all, if Tesla can get away with it, only suffering the occasional wrist slap for egregious rule violations, is it not the case that other companies have a duty to their stockholders to do the same so as to compete on a level playing field? As long as we (the public) allow companies to get away with blaming drivers for crashes instead of pinning responsibility on auto makers for deploying poorly executed automation features, reckless road testing (for example, running red lights) will continue to proliferate.

While it is difficult to draw a line in the sand, we can propose one. Any vehicle that can make turns at intersections with no driver control input is not a Level 2 system -- it is an autonomous vehicle prototype. There is an argument supporting this based on the evolution of J3016. (See page 242 of this paper.) But more importantly, supervising turns in traffic is clearly of a different nature than supervising cruise control on a highway. It takes much more attention and awareness of what the vehicle is about to do to prevent a tragic crash. Especially if you are busy clapping hands to a rock song.

Abusing Level 2/2+ to evade regulation

Another trail blazed by Tesla is letting regulators think they are deploying unregulated Level 2/2+ systems when they are really putting Level 3/4 testbeds on the roads in the hands of ordinary vehicle owners.

Companies should not call something Level 2+ when it is really a testbed for Level 3 features.  Deployed Level 2+ features must be production ready instead of experimental, and should never put human drivers in the role of being an unqualified "tester" of safety critical functionality. Among other things, the autonowashing label of "beta test" that implies drivers are responsible for crashes due to automation malfunctions should be banned. 

The details are subtle here: if a vehicle does something dangerous that is not reasonably expected by an ordinary driver, that should be considered a software malfunction. Any dangerous behavior that cannot be compensated for by an ordinary driver should not happen -- regardless of whether the driver has been told to expect it or not. Products sold to retail customers should not malfunction, nor should they exhibit unreasonably dangerous behaviors. Telling the driver they are responsible for safety should not change the situation.

Software should either work properly or be limited to trained testers operating under a Safety Management System. Rollout of Level 2+ and Level 3 features can happen in parallel, and that's fine. But a prototype Level 3 feature that can malfunction is not in any way the same as a mature Level 2+ feature -- even if they superficially behave the same way. It's not about the intended behavior, it's about having potentially very ill-behaved malfunctions. A Level 2+ system might not "see" a weird object in the road and perhaps the driver should be able to handle such a situation if they have been told that is an expected behavior. But any system that suddenly swerves the vehicle left into oncoming traffic is not fit for use by everyday drivers. A suddenly left-swerving vehicle is not Level 2/2+ -- it is a malfunctioning prototype vehicle that should only be operated by trained test drivers regardless of the claimed automation level.

Over time we'll see car companies aggressively try to get features to market. We can expect repeated iterations of arguments that the next incremental feature is no big deal for a driver to supervise. But if that feature comes with cautions that drivers must pay special attention beyond what a driver monitoring system supports, or incur extra responsibility, that should be a red flag that we're looking at asking them to be testers rather than drivers. Warnings that the vehicle "may do the wrong thing at the worst time" are unambiguously a problem. Such testing should be strictly forbidden for deployment to ordinary drivers. 

Telling someone that a product is likely to have defects (that's what "beta" means these days, right?) means that they are being sold a likely-defective product. As such, they should not be distributed as retail products to everyday customers, regardless of click-through disclaimers. After all, any unqualified "tester" is not just placing themselves at risk, but other road users as well.

For those who say that this will impede progress, see my discussion of the myths propagated by the AV industry in support of avoiding safety regulation. Following basic road testing safety principles should be a part of any development process. Anything else is irresponsible and, in the long run, will hurt the industry more than it helps by generating bad press and accompanying ill will.

Civilian drivers should never be held responsible for malfunctioning technology by being used as Moral Crumple Zones. That especially includes mischaracterizing Level 3 (and above) prototype features as Level 2+ systems.

The bottom line 

Perhaps NHTSA will wake up and put a regulatory framework in place to help the industry succeed at autonomous vehicle deployment safety. (They already have proposed a plan, but it is in suspended animation.) Even if the new NHTSA framework proceeds, regulations and proposals from both NHTSA and the states address levels 3-5, while leaving a situation ripe for Level 2+ abuse. Until that changes, expect to see companies pushing the envelope on Level 2+ as they compete for dominance in the vehicle automation market.

Saturday, November 27, 2021

Regulating Automated Vehicles with Human Drivers

Summary 

Regulatory oversight of automated vehicle operation on public roads is being gamed by the vehicle automation industry via two approaches: (1) promoting SAE J3016, which is explicitly not a safety standard, as the basis for safety regulation, and (2) using the "Level 2 loophole" to deploy autonomous test platforms while evading regulatory oversight. Regulators are coming to understand they need to do something to reign in the reckless driving and other safety issues that are putting their constituents at risk. We propose a regulatory approach to deal with this situation that involves a clear distinction between production "cruise control" style automation that can be subject to conventional regulatory oversight vs. test platforms that should be regulated via SAE J3018 use for testing operational safety.

Video showing Tesla FSD beta tester unsafely turning into oncoming traffic.

Do not use SAE J3016 in regulations

The SAE J3016 standards document has been promoted by the automotive industry for use in regulations, and in fact is the basis for regulations and policies at the US federal, state, and municipal levels. However, it is fundamentally unsuitable for the job. The issues with using SAE J3016 for regulations are many, so we provide a brief summary. (More detail can be found in Section V.A of our SSRN paper.)

SAE J3016 contains two different types of information. The first is a definition of terminology for automated vehicles, which is not really the problem, and in general could be suitable for regulatory use. The second is a definition of the infamous SAE Levels, which are highly problematic for at least the following reasons:

In practice, one of two big issues is the "Level 2 Loophole" in which a company might claim that the fact there is a safety driver makes its system Level 2, while insisting it does not intend to ever release that same automated driving feature as a higher level feature. This could be readily gamed by, for example, saying that Feature X, which is in fact a prototype fully automated driving system, is Level 2 at first. When the company feels that the prototype is fully mature, it could simply rebrand it Feature Y, slap on a Level 4 designation, and proceed to sell that feature without ever having applied for a Level 4 testing permit. We argue that this is essentially what Tesla is doing with its FSD "beta" program that has, among other things, yielded numerous social media videos of reckless driving despite claims that its elite "beta test" drivers are selected to be safe (e.g., failure to stop at stop signs, failure to stop at red traffic signals, driving in opposing direction traffic lanes).

The second big practical issue is that J3016 is not intended to be a safety standard, but is being used as such in regulations. This is making regulations more complex than they need to be, stretching the limits of the esoteric technical expertise in AVs required of regulatory agencies, especially for municipalities. This is combined with the AV industry promoting a series of myths as part of a campaign to deter regulator effectiveness at protecting constituents from potential safety issues. The net result is that most regulations do not actually address the core safety issues related to on-road testing of this immature technology, in large part because they aren't really sure how to do that.

For road testing safety purposes, regulators should focus on both the operational concept and technology maturity of the vehicle being operated rather than on what might eventually be built as a product. In other words "design intent" isn't relevant to the risk being presented to road users when a test vehicle veers into opposing traffic. Avoiding crashes is the goal, not parsing overly-complex engineering taxonomies.

The solution is to reject SAE J3016 levels as a basis for regulation, instead favoring other industry standards that are actually intended to be relevant to safety. (Again, using J3016 for terminology is OK if the terms are relevant, but not the level definitions.)

Four Regulatory Categories

We propose four regulatory categories, with details to follow:

  • Non-automated vehicles: These are vehicles that DO NOT control steering on a sustained basis in any operational mode. They might have adaptive speed control, automatic emergency braking, and active safety features that temporarily control steering (e.g., an emergency swerve around obstacles capability, or bumping the steering wheel at lane boundaries to alert the driver).
  • Low automation vehicles: These are vehicles with automation that CAN control steering on a sustained basis (and, in practice, also vehicle speed). They are vehicles that ordinary drivers can operate safely and intuitively along the lines of a "cruise control" system that performs lane keeping in addition to speed control. In particular, they have these characteristics:
    • Can be driven with acceptable safety by an ordinary licensed driver with no special training beyond that required for a non-automated version of the same vehicle type.
    • Includes an effective driver monitoring system (DMS) to ensure adequate driver alertness despite inevitable automation complacency
    • Deters reasonably foreseeable misuse and abuse, especially with regard to DMS and its operational design domain (ODD)
    • Safety-relevant behavioral inadequacies consist of omissive behaviors rather than actively dangerous behavior
    • Safety-relevant issues are both intuitively understood and readily mitigated by driver intervention with conventional vehicle controls (steering wheel, brake pedal)
    • Automation is not capable of executing turns at intersections.
    • Field data monitoring indicates that vehicles remain at least as safe as non-automated vehicles that incorporate comparable active safety features over the vehicle life.
  • Highly automated vehicles: These are vehicles in which a human driver has no responsibility for safe driving. If any person inside the vehicle (or a tele-operator) can be blamed for a driving mishap, it is not a highly automated vehicle. Put simply, it's safe for anyone to go to sleep in these vehicles (including no requirement for a continuous remote safety driver) when in automated operation.
  • Automation test platforms: These are vehicles that have automated steering capability and have a person responsible for driving safety, but do meet one or more of the listed requirements for low automation vehicles. In practical terms, such vehicles tend to be test platforms for capabilities that might someday be highly automated vehicles, but require a human test driver -- either in vehicle or remote -- for operational safety.
Non-automated vehicles can be subject to regulatory requirements for conventional vehicles, and correspond to SAE Levels 0 and 1. We discuss each of the remaining three categories in turn.

Low automation vehicles

The idea of the low automation vehicle is that it is a tame enough version of automation that any licensed driver should be able to handle it. Think of it as "cruise control" that works for both steering and speed. It keeps the car moving down the road, but is quite stupid about what is going on around the car. DMS and ODD enforcement along with mitigation of misuse and abuse are required for operational safety. Required driver training should be no more than trivial familiarization with controls that one would expect, for example, during a car rental transaction at an airport rental lot.

Safety relevant issues should be omissive (vehicle fails to do something) rather than errors of commission (vehicle does the wrong thing). For example, a vehicle might gradually drift out of lane while warning the driver it has lost lane lock, but it should not aggressively turn across a centerline into oncoming traffic. With very low capability automation this should be straightforward (although still technically challenging), because the vehicle isn't trying to do more than drive within its lane. As capabilities increase, this becomes more difficult to design, but dealing with that is up to the companies who want to increase capabilities. We draw a hard line at capability to execute turns at intersections, which is clearly an attempt at high automation capabilities, and is well beyond the spirit of a "cruise control" type system.

An important principle is that human drivers of a production low automation vehicle should not serve as Moral Crumple Zones by being asked to perform beyond civilian driver capabilities to compensate for system shortcomings and work-in-progress system defects. If human drivers are being blamed for failure to compensate for behavior that would be considered defective in a non-automated vehicle (such attempting to turning across opposing traffic for no reason), this is a sign that the vehicle is really a test platform in disguise.

Low automation vehicles could be regulated by holding the vehicles accountable to the same regulations as non-automated vehicles as is done today for Level 2 vehicles. However, the regulatory change would be excluding some vehicles currently called "Level 2" from this category if they don't meet all the listed requirements. In other words, any vehicle not meeting all the listed requirements would require special regulatory handling.

Highly automated vehicles

These are highly automated vehicles for which the driver is not responsible for safety, generally corresponding to SAE Levels 4 and 5.  (As a practical matter, some vehicles that are advertised as Level 3 will end up in this category in practice if they do not hold the driver accountable for crashes when automation is engaged.)

Highly automated vehicles should be regulated by requiring conformance to industry safety standards such as ISO 26262, ISO 21448, and ANSI/UL 4600. This is an approach NHTSA has already proposed, so we recommend states and municipalities simply track that topic for the time being. 

There is a separate issue of how to regulate vehicle testing of these vehicles without a safety driver, but that issue is beyond the scope of this essay. 

Automation test platforms

These are vehicles that need skilled test drivers or remote safety monitoring to operate safely on public roads.  Operation of such vehicles should be done in accordance with SAE J3018, which covers safety driver skills and operational safety procedures, and should also be done under the oversight of a suitable Safety Management System (SMS) such as one based on the AVSC SMS guidelines.

Crashes while automation is turned on are generally attributed to a failure of the safety driver to cope with dangerous vehicle behavior, with dangerous behavior being an expectation for any test platform. (The point of a test platform is to see if there are any defects, which means defects must be expected to manifest during testing.)

In other words, with an automation test platform, safety responsibility primarily rests with the safety driver and test support team, not the automation. Test organizations should convince regulators that testing will overall present an acceptably low risk to other road users. Among other things, this will require that safety drivers be specifically trained to handle the risks of testing, which differ significantly from the risks of normal driving. For example, use of retail car customers who have had no special training per the requirements of SAE J3018 and who are conducting testing without the benefit of an appropriate SMS framework should be considered unreasonably risky.

This category covers all vehicles currently said to be Level 4/5 test vehicles, and also any other Level 2 or Level 3 vehicles that make demands on driver attention and reaction capabilities that are excessive for drivers without special tester training.

Regulating automated test platforms should concentrate on driver safety, per my State/Municipal DOT regulatory playbook. This includes specifically requiring compliance with practices in SAE J3018 and having an SMS that is at least as strong as the one discussed in the AVSC SMS guidelines.

Wrap-up

Automated vehicles regulatory data reporting at the municipal and state levels should concentrate on collecting mishap data to ensure that the driver+vehicle combination is acceptably safe. A high rate of crashes indicates that either the drivers aren't trained well enough, or the vehicle is defective. Which way you look at it depends on whether you're a state/municipal government or the US government, and whether the vehicle is a test platform or not. But the reality is that if drivers have trouble driving the vehicles, you need to do something to fix that situation before there is a severe injury or fatality on your watch.

The content in this essay is an informal summary of the content in Section V of: Widen, W. & Koopman, P., "Autonomous Vehicle Regulation and Trust" SSRN, Nov. 22, 2021. In case of doubt or ambiguity, that SSRN publication should be consulted for more comprehensive treatment.

-----
Philip Koopman is an associate professor at Carnegie Mellon University specializing in autonomous vehicle safety. He is on the voting committees for the industry standards mentioned. Regulators are welcome to contact him for support.


Sunday, September 12, 2021

Autonomous Vehicle Testing Guidance for State & City DOTs

Once in a while I'm contacted by a city or state Department of Transportation (DOT) to provide advice on safety for "self-driving" car testing. (Generally that means public road testing of SAE Level 3-5 vehicles that are intended for eventual deployment as automated or autonomous capable vehicles,)

The good news is that industry standards are maturing. Rather than having to create their own guidelines and requirements as they have in the past, DOTs now have the option of primarily relying upon having AV testers conform to industry-created guidelines and consensus standards.

And ... in September 2021 NYC DOT blazed a trail by requiring the self-driving car industry to conform to their own industry consensus testing safety standard (J3018). Kudos to NYC DOT!  (check it out here (link); more on that in the details below.

The #1 important thing to keep in mind is that testing safety is not about the automation technology -- it is about the ability of the human safety driver to monitor and intervene when needed to make safety. The technology is going to fail, because the point of testing is to find surprise failures. If a failure of technology causes a fatality, then most likely the testing wasn't being done safely enough. It is essential that human safety drivers be skilled and attentive enough to prevent loss events when such failures inevitably occur.

The short version is that DOTs should:

  1. Follow the AAMVA road testing guidelines plus some additional key practices.
  2. Define how safe testing should be when considering the safety driver + vehicle system as a whole.
  3. Ask testers for conformance to SAE J3018 for road testing.
  4. Ask testers to have a credible Safety Management System (SMS) approach, including a testing plan.
  5. Ask testers to provide metrics that show that their testing is safe (not just a promise up front, but also periodic testing safety metrics as they operate).  Don't get distracted by measuring the maturity of the technology they are testing -- it's all about the safety driver ability to intervene when something goes wrong.
  6. If testing takes place with a safety driver in a chase vehicle or remote, ask for conformance to safety standards for the mechanisms required to ensure safety (e.g., per ISO 26262), but otherwise conforming to SAE J3018 for the training and protocols.
  7. If testing takes place without continuously monitoring safety driver, ask for conformance to industry-consensus safety standards for the autonomous vehicle itself.  If there is no person continuously monitoring and capable of assuring safety, then the safety aspects of the technology have to be done. You shouldn't let vehicles without fully mature safety technology operate without a human safety driver.


Cars driving from cell phone to road


The long version (below) gets pretty detailed, but this is a complicated and nuanced issue. So, here we go...

The AAMVA Guidelines as a starting point:

DOTs should follow applicable AAMVA guidelines with a few additional points.

The American Association of Motor Vehicle Administrators (AAMVA) released the 2nd edition of Safe Testing and Deployment of Vehicles Equipped with Automated Driving Systems Guidelines in September 2020. There is plenty of good information here. However, there are a few areas that require going beyond these guidelines to ensure what might be considered acceptable safety.

My additional recommendations within the scope of these guidelines include:
  • Vehicle manufacturer or testing organization should be required to publish a Voluntary Safety Self Assessment (VSSA) report (see AAMVA Guidelines 3.1.5). That VSSA should address all relevant topics in the NHTSA Automated Driving Safety documents (2.0, 3.0, 4.0). A VSSA does not provide all information required for technical evaluation of safety, but complete lack of a VSSA suggests an unwillingness to provide public transparency.
  • Require statement of areas of intended operation in a manner that does not compromise any claimed secrets as to detailed specifics of tests being conducted. 
    • For example, require reporting of zip codes of where testing is to be conducted.
    • Report speeds at which testing will be conducted (e.g., 25 mph speed limit street testing is much different than Interstate System highway testing).
    • Report other relevant Operational Design Domain factors that will limit testing (e.g., daytime only, in rain, in snow) so that any particularly hazardous environmental testing situations can be considered with regard to public safety.
    • Discuss any unique characteristics of the test area to ensure the tester understands what unique challenges might be presented that someone not from the location might find unusual (e.g., The Pittsburgh Left, parking chairs, cable cars, cattle grids, gator crossings).
  • Require a tester statement that a defined level of technology quality will be confirmed before it is used for public road testing (along with the definition of what that might be). This should include at least:
    • A comprehensive simulation and closed course testing plan should be completed before testing on public roads.
    • All software updates should be subjected to confirmatory closed course testing to ensure no new defects have been introduced before being used in road testing.
    • Any vehicle feature that does not pass closed course testing should not be active during road testing. (In other words, if a feature fails closed course track testing, it shouldn't be operated on public roads.) Public roads should be used to confirm that the vehicle works as expected, not for debugging of known-faulty features.
  • Explanation for why the tester thinks that safety driver training and performance will be sufficient to ensure that test vehicles do not present increased risk to other road users.
The AAMVA guideline scope, while quite useful, is primarily administrative in nature rather than technical. To go beyond this we need to look at engineering standards. (Some of the above points also appear in the following standards and guidance.)

Define how safe is safe enough:

DOTs should define the desired safety outcome, but not how to measure it.

This is perhaps the trickiest point. It's important for the DOT to set the bar for how safe is safe enough. Testers likely have overwhelming financial incentive to get their testing done. Even with the best intentions, the threat of losing funding for lack of progress can loom larger than a possibility of a problem with a testing crash that might (or might not) happen in the future. It seems insufficient in such an environment to simply assume that for-profit organizations will set a safety target that reflects local societal norms.

However, it would be irresponsible for a testing organization to do public road testing without regard for public safety. This means that any (responsible) testing organization will have: a safety goal and analysis before testing starts to predict whether they are likely to reach that safety goal, and metrics collected during testing to ensure that they are meeting their safety goal.

DOTs might not have the technical sophistication to tell testers how to predict safety during testing, nor to know exactly which metrics and associated metric thresholds would be appropriate for a particular test plan. However, the DOT should take responsibility (absent legislation) for making it clear what the testing safety goal should be.

An example might be: road testing operations shall be at least as safe as unimpaired human drivers, taking into account local driving safety statistics and testing environmental conditions. For example, if testing in Pittsburgh only in daytime and dry weather, testers should have a goal of being at least as safe as other Pittsburgh drivers operating in daytime and dry weather (subtracting out drunk and impaired human driver collisions).  That "safer than human" should consider at least fatalities and major injury crashes. Records must be kept of all safety-related metrics, incidents, and loss events.

Some important considerations is that the policy in the preceding paragraph does not tell testers how to predict such safety nor how to measure it on a technical basis. Rather, it is up to the testers to figure this out in their own individual situation. As mentioned earlier, if they don't know how to measure their own safety, they shouldn't be out on public roads doing the testing in the first place.

Could this approach be gamed? Of course it can (as can any approach). However, if the tester goes on record committing to a particular level of safety, it will become evident whether that level of safety has been reached sooner or later based on police reports, if nothing else. Once that happens, historical metrics will show whether the tester was operating in good faith or not.

SAE J3018 for operational safety:

DOTs should ask testers to conform to the industry standard for road testing safety: SAE J3018. 

When AV road testing first started, it was common for testers to claim that they were safe because they had a "safety driver." However, as was tragically demonstrated in the Tempe AZ testing fatality in 2018, not all approaches to safety driving are created equal.  Much more is required. Fortunately, there is an SAE standard that addresses this topic.

SAE J3018_202012 "Safety-Relevant Guidance for On-Road Testing of Prototype Automated Driving System (ADS)-Operated Vehicles" (https://www.sae.org/standards/content/j3018_202012/ -- be sure to get the 2020 revision) provides safety relevant guidance for road testing. It concentrates on guidance for the "in-vehicle fallback test driver" (also known informally as the safety driver).

AV testers should be conform to J3018 to ensure that they are following identified best practices for safety driver training and effectiveness. Endorsing this standard will avoid a DOT having to create their own driver qualification and training requirements.

Taking a deeper look at J3018, it seems a bit light on measuring whether the safety driver is actually providing effective risk mitigation. Rather, it seems to implicitly assume that training will necessarily result in acceptable road testing safety. While training and qualification of safety drivers is essential, it is prudent to also monitor safety driver effectiveness, and testers should be asked to address this issue. Nonetheless, J3018 is an excellent starting point for testing safety. Testers should be doing at least what is in J3018, and probably more.

J3018 does cost money to read, and the free preview is not particularly informative. However, there is a free copy of a precursor document available here: https://avsc.sae-itc.org/principles-01-5471WV-42925L3.html  that will give a flavor of what is involved. That having been said, any DOT guidance or requirement should follow J3018, and not the AVSC precursor document.

In addition to following J3018, the safety-critical mechanisms for testing should be designed to conform to the widely used ISO 26262 functional safety standard. (This is not to say that the entire test vehicle -- which is still a work in progress -- needs to conform to 26262 during testing. Rather, that the "Big Red Button" and any driver takeover functions need to conform to 26262 to make sure that the safety driver can really take over when necessary.)

For cargo vehicles that will deploy without drivers, J3018 can still be used by installing a temporary safety driver seat in the vehicle. Or the autonomy equipment can be mounted on a conventional vehicle in a geometry that mimics the cargo vehicle geometry. When the time comes to deploy without a driver physically in the system, you are really testing an autonomous vehicle with a chase car or remote safety supervisor, covered in a following section on testing without a driver.

Safety Management System (SMS)

DOTs should ask testers to have a Safety Management System in place before testing.

A Safety Management System is a systematic way to manage safety risk for an organization. The roots of SMS approaches come from the aviation industry. The short version is that an SMS helps make sure that you are operationally safe. An important aspect of an SMS is that traditionally it is more about how the people in the company perform tasks and the safety culture rather than the technology itself.

Perhaps the most important overarching finding of the NTSB investigation of the Tempe AV testing fatality was that the lack of an SMS increased the risk of such a bad outcome. To paraphrase the NTSB hearing opening remarks: "you don't have to wait to have a fatal crash before you decide to implement an SMS."  (If you have made it this far in reading this essay, you absolutely must listen to the first 6 minutes of this NTSB hearing https://youtu.be/mSC4Fr3wf0k if you have not already done so.)

The AVSC, a closed-membership industry group, has recently released guidelines for AV testing SMS: https://avsc.sae-itc.org/principle-7-5896VG-46559OG.html  
While these are not at the same level of consensus and review of an SAE issued standard (for example, public comments are not solicited), they do provide industry guidance that is applicable to road testing safety. (I personally have not reviewed these to the degree I have J3018, but expect to do so over time if it is submitted to the SAE ORAD standards committee as J3016 was. So this is not a specific endorsement, but rather an identification of industry-created content that looks likely to be useful.)

DOTs should ask that any AV testing organization to describe their SMS and accompanying safety plan. The tester should explain how such an SMS is comparable to or better than the AVSC guidelines.

Metrics:

DOTs should ask for metrics related to public safety during testing rather than autonomy performance.

It is common to want a standard set of metrics for both test and deployed vehicles. That area is still maturing. While metrics such as number of crashes of various severity classes are fairly straightforward, other predictive metrics such as "disengagements" are problematic for a number of reasons. In particular, each vehicle and each test program has different objectives and different safety architectures. So it will be a while before one-size-fits-all metrics are standardized.

We recommend that any metrics defined be tied to safety procedures and policies rather than the maturity of the technology. Most importantly, it is desirable to find metrics that AV testers cannot claim reveal proprietary information. That means that metrics that measure "how good is the AV" or "how soon to deployment" are likely to be problematic -- and not necessarily that relevant to the crucial question of whether the testing itself that's going to happen right now (and not the AV that might be deployed sometime in the future) presents elevated risk to the public.

I'd argue that the public has a legitimate right to understand whether road users are in the test area are put at increased risk due to AV testing. One way to approach this is to ask the AV tester to respond the following questions:
  • What basis do you have for claiming that your testing will not present increased risk to other road users, including vulnerable road users?
  • What metrics to you plan to collect to ensure that your system is in fact not presenting any such increased risk?
  • What periodic (e.g., monthly) quantitative report can you give us to show that indeed your testing has not increased the risk to other road users?
In general, the strategy should be to ask the AV tester: "Why do you think you're safe" and "How do you plan to measure safety," followed by "How will you know if you're not as safe as you promised you would be?"

If the AV testing can't promise that they will not increase risk to other road users (especially vulnerable road users), then should they be testing on your roads?  If they don't plan to measure and track their actual on-road risk, then do you find their safe testing promise credible? And if they claim that their testing road risk data is proprietary, does that even make sense?

Some example metrics for testing safety (although applicability depends on the specifics of the situation):
  • How often does the built-in driver monitor signal a driver attention issue?  (It won't be zero, but there should be a defined acceptable threshold set by the AV tester which, if exceeded, should cause a process intervention of some sort.)
  • How often does the safety driver make an erroneous intervention, even though there is no crash or other loss event? (In other words, how many near hits are occurring?)
  • How does the AV tester track skill degradation to determine when it is time for a shift change or even refresher training for a safety driver?
Keep in mind that for most companies testing safety is accomplished via test driver supervision, road safety has much more to do with the reliability of the safety drivers than the automation technology itself. So the above metrics have nothing to do with the automation technology, and everything to do with test driver safety -- which is the part that matters for most AV testing safety.

In the end, the metrics should show that the required level of safety is being achieved. They should also be predictive enough that they are likely to indicate any potential problems BEFORE there is a crash.

Testing Without A Driver:

DOTs should ask about safety during communication loss for remote safety driver testing.
DOTs should ask testers to conform to industry automotive safety standards if there is no supervising test driver.

Eventually, organizations will want to test on public roads without a driver. Indeed California has already issued permits for this. In terms of safety, a primary question to ask is how safety is being assured. 

If there is a remote operator involved, then it is important to ensure that any real time data connectivity and sensor information is sufficient to ensure safety. This is a controversial area, and any company promising that, for example, a remote operator can instantly take over operation in the event of a malfunction should be prepared to offer hard data metrics on control latency (delay introduced by the remote communication system), effectiveness of the vehicle detecting its own malfunctions (very difficult to ensure if the system doesn't know it doesn't see a pedestrian for example), and communication link reliability. It is challenging (some would say implausible) to control high speed vehicle operation remotely due to the latencies involved, so a line of sight radio link with a chase car might be required. J3018 practices for the remote operator would still apply. Additionally the equipment used to perform the remote operation should conform to ISO 26262 or other comparable safety standard, which is not typically true for telecommunication equipment. (If loss of signal triggers a vehicle shutdown, then that loss of signal equipment and shutdown mechanism should conform to ISO 26262.)

If there is no remote operator involved, then either the tester should be following issued safety standards or have a safety case to explain why what they are doing is at least as rigorous as what is in those standards. Currently issued and applicable safety standards include: ISO 26262 (functional safety), ISO 21448 (safety of the intended function), and ANSI/UL 4600 (system level safety for autonomous vehicles).

It is worth noting that misinformation has been provided to at least one state DOT regarding ANSI/UL 4600 by industry advocacy groups. (Short version: there is no requirement whatsoever for external assessment in 4600, despite multiple statements to the contrary in a letter sent to a state DOT. Other negative statements tend to be similarly misleading or just plain incorrect.) Any DOT who wants the full story in response to any information they receive criticizing ANSI/UL 4600 is welcome to contact the author of this essay.

Some testers may say they have reasons for not following industry consensus safety standards. If that is the case, ask them what quantitative data they have to demonstrate they are safer than a human driver. If they can't prove to themselves that they are at least as safe as a human driver, why are they operating on public roads? If they say they have the data but it is proprietary, ask what road testing safety data has to do with the secret sauce behind their autonomy.  (Short answer -- it has nothing to do with the secret sauce, but might have to do with concerns that they can't promise safe testing.)

Transparency:

It is common for testers to claim that any attempt to require data reporting, metrics, or other transparency will somehow give away incredibly valuable trade secrets and inhibit innovation. This is utter nonsense. Yet, it seems to be the industry playbook.  For example, during a NYC DOT hearing "about a half-dozen autonomous car makers and their advocates said the proposed rules would turn New York City from an engine of innovation into a backwater that would set back the evolution of the potentially life-saving technology of computer-controlled cars and trucks that can move around without inferior human beings messing everything up."  (https://nyc.streetsblog.org/2021/09/01/self-driving-car-industry-promising-safety-pushes-back-on-dot-plan-to-regulate-testing/)

Often this conversation boils down to testers saying "trust us, we're smart." They may be smart, but decades of experience with safety in other domains has shown that there is no safety without transparency. If they are smart enough to be able to build a car that can drive itself safely on your roads -- without even needing to follow industry standards -- they should also be smart enough to figure out a way to show you data to prove they are safe without revealing major secrets.

For situations in which a safety driver is in the vehicle, let's look at what is required for transparency, which NYC DOT did a good job with (here: https://rules.cityofnewyork.us/wp-content/uploads/2021/08/DOT-Notice-of-Adoption-AV-Rule-FINAL-with-Finding.pdf). The elements they require are:
  • Self-certification that the testing will be safer than a human driver. This is just asking the tester to claim (without producing any proof) that they will test safely. If they're not willing to sign up to that, probably they should not be on public roads.
  • Conform to SAE J3018 and AVSC 00001201911. In other words, this is asking them to follow industry standards and practices for their test driver qualification and testing protocols. This involves ONLY the human test driver and does not place constraints on the automation technology being tested. If they're not willing to sign up to have trained safety drivers and safe testing protocols, probably they should not be on public roads.
  • Submission of a safety plan. This has nothing to do with the automation technology -- it is all about making sure the safety driver can keep the vehicle safe. If they can't explain to the DOT what their plan is to be safe in testing, probably they should not be on public roads.
The key is: you don't need to disclose any autonomous vehicle secret sauce to explain why testing will be safe, because the safety hinges on the human safety driver, not the automation technology.

Other Resources.

Here are some resources that might be useful. While SAE J3016 is widely used for terminology, it is essential to note that it is not (and is not intended to be) a safety standard. Conformance to J3016 Levels has to do with whether you're using an appropriate name for your vehicles, and not whether those vehicles are safe. 

Friday, December 4, 2020

Road Testing Safety Metrics (Metrics Episode 3)

To build trust, self-driving car companies should be transparent about operational safety metrics for road testing, such as how often the safety driver fails to react to an issue. This won't be perfect, but it ought to be at least as safe as normal human drivers on public roads.

Right now when you see what looks like a self driving car on the road, it’s not really a production self driving car -- it’s a self driving car technology test platform. That usually means that some human someplace in the car remotely, whatever, is keeping an eye on it and making sure things are safe. In turn if you asked the question, “How safe is that car,” you’re not actually asking about the safety of the self driving technology pretty much at all. What you’re asking is whether or not that human is able to properly supervise the safety and do the right thing when something goes wrong. So if you care about the safety of this public on road testing of this maturing technology what you care about is the human safety driver’s performance. To understand that let’s talk about the types of things that safety driver has to do.

The safety driver has to build and maintain situational awareness and know what’s supposed to happen next. The safety driver has to notice something is going wrong and intervene at exactly the right time. While this is happening the safety driver is under some pressure to balance getting operational data with maintaining safety. After all, sitting in the garage with the car turned off doesn’t get the data that they’re out in the public roads to get. As they’re operating the safety driver has to figure out how to react when other drivers do things that are weird, illegal, dangerous, or just plain crazy.  And sometimes it’s the car itself that misbehaves due to a software defect in the self driving technology. When something bad happens, the safety driver has to execute a takeover maneuver and make sure that they don’t make things worse.

Now, that sounds like a lot to do but the fact is that when things are going well it’s an exceptionally boring job. The car just keeps doing the thing it’s supposed to be doing and the driver sits there watching and waiting. Every once in a while something will go wrong and the safety driver has to make sure that not only do they have situational awareness, and that they have an idea of what to do next, but also that they’re not startled into doing the wrong thing. Rolling all that up, supervising an autonomous test platform is exceptionally difficult work and requires the highest of driving skills. And yes, we actually do get to the part where the car has to behave properly, but it has nothing to do with the self driving technology -- it has to do with the emergency override. If the human safety driver wants to override the vehicle then they have to be able to turn off the self driving feature and get control of the vehicle.

It turns out getting that mostly right is straight forward but getting it 100% right takes a lot of care and attention. Taking all this into account, you start realizing that disengagements, which is the measure we use now, is actually the wrong metric. In fact it’s exactly the wrong metric. That's because disengagements aren’t what make you dangerous, the thing that makes you the most dangerous when you’re testing technology is non disengagements: a failure to disengage when you should have. In other words, the hazards that you did not catch with the human safety driver that’s what matters.

Now, measuring something that doesn’t happen is not straightforward but in fact if you want metrics for safety that’s what you have to go after. The starting point isn’t really quite metrics at all but rather processes and what you hear is the more thoughtful companies are starting to emphasize things like good driver qualifications, they have training, they have good reaction times. Driver testing, go out on a test track, make sure the driver can handle surprises and make sure the driver can maintain situational awareness. Yes, the disengagement mechanism has to work. The proverbial big red button needs to work all the time not just most of the time. Fortunately there’s a safety standard ISO 26262 which tells you exactly how to make that kind of system so these companies should be following that standard with their disengagement mechanism.

Now there’s an issue about whether or not training once is enough.  It’s not. You need refreshers at least, and so refresher training is also a good idea. Many of the companies stop with that. But the problem with doing only training is that you think you’re good enough but you don’t know you’re actually good enough unless you take operational data. 

Sure, the driver can pay attention on test day and maintain engagement. But can a safety driver maintain that same peak performance for six months of day in, day out testing that’s really boring because mostly things work pretty well? Humans aren’t perfect, and that means even really good safety drivers aren’t perfect.  That’s just the way it is. So what you really want is not just training but also operational metrics to make sure that the safety drivers are able to be as good as they need to be to reach your safety goal.

Here are some metrics that could work internally for an engineering effort:

  • You could measure the attention of the driver throughout a shift. Maybe you’ll find out, as might be expected, that after many hours the driver has trouble maintaining attention. That wouldn’t be a surprise. Some companies are doing two hour shifts. While for active driving that might be a good shift size, there’s lots of data for several decades suggesting that for supervising autonomy you need shorter shifts, maybe only 30 minutes. So do you have data showing that your two hour shifts are okay? Or maybe you need a shorter shift?  Without data you won’t know that.
  • You might also want to measure the intervention accuracy. Did the driver actually know if their blind spot was clear when they made a lane change, or did they just get lucky when they did a takeover and a lane change? Knowing the fraction of time that drivers do the right thing is important. It’s not going to be 100%, but you just need to make sure it’s good enough to achieve your safety goal. 
  • And every once in a while if the driver tries to do a takeover and the vehicle doesn’t respond you really need to know that that happened.

Now those metrics are somewhat detailed. It might not be the kind of thing that a government agency or the public would really be in a position to process well. So you want some roll up metrics as well. But these are mostly things that have to do with non disengagements rather than disengagements. 

To really deal with road testing safety, the number one concern has got to be times when the vehicle did something dangerous and the safety driver did not do the right compensating action. Your drivers aren’t going to be perfect. Are you planning on just blaming them when there’s a mishap or are you actually taking measurements to find out if you’re hitting performance targets?  So if you have a target for how much the driver has to pay attention and how much is okay to lose attention for a second or two, then you need to know if you’re hitting that target and it’s not going to be perfection.

You probably also want some metrics for rule violations. 

  • For example, running a stop sign even though nothing bad happened, is still a bad thing. The number of times that happens is not going to be zero, but it should be so low that it’s dramatically better than normal human drivers for example.
  • You probably want to measure how often you unreasonably encroach on a bike lane even if there’s no bicyclists there.
  • You might want to measure how often you fail to yield to a pedestrian. Most pedestrians aren’t going to jump out in front of a car to get hit, but the fact that they had to back off and let you pass means you weren’t behaving the way you were supposed to. 
  • Probably the most important thing to measure is near misses (or as some people call them near hits) where the margin was just too small even if you got lucky and nothing bad actually happened. For example, maybe you’re supposed to leave three feet of clearance to a bicyclist. (And at high speeds I would hope you do leave at least that, but if you’re down at two or three miles an hour crawling through a dense urban area and you're two feet away from a bicyclist, maybe that’s okay.) You need a definition of near misses that responds to the particular situation but ultimately you want to know how often you had a near miss, meaning you were in a situation more dangerous than it was supposed to be. If your safety driver catches it, then great. But if your safety driver does not, that means you’re operating taking chances that you didn’t intend to take, and knowing that is super important.

No safety program and no safety metric is going to be perfect. It’s unreasonable to expect self driving car testing to be absolutely perfect with not even a fender bender. Surely mishaps will happen especially when interacting with human drivers. But there has to be some sort of strategy. 

A viable strategy might be that you use above average safety drivers who have to deal with the extra complexity of the safety platforms, but they come out overall to be at least as good as normal human drivers. You’re making a big ask to have these safety drivers be better than normal human drivers and compensate for all the difficulties with this technology. Probably it can be done, but you really need to know you got there instead of training the drivers, sending them out on roads and simply hoping that you’re safe enough. You really want some metrics. 

To build trust companies should be exposing at least the high level rollup metrics to the public. They have nothing to do with the secret sauce behind the self driving car but they have everything to do with the safety of the general public as this technology is being tested on public roads.