Process safety management is a management system for preventing the release of energy or hazardous material from a process. It is not a document set, and it is not a department. It is the arrangement by which a facility knows what it is handling, understands how that can hurt people, keeps the plant inside the envelope where the understanding still holds, and learns when it does not.
Almost every serious process incident of the last forty years was preceded by a functioning management system that had drifted in one or two specific places. That is why the frameworks are organised by element rather than by hazard. The elements are the places drift is known to occur, and the point of naming them is so that each one can be examined on its own evidence.
This piece sets out the element sets, what each element has to produce, the thresholds that decide whether the regime applies at all, and the arithmetic behind the indicators used to judge performance. It is written for engineers and managers who have to build, assess or defend a system rather than for readers who want the idea.
Why there is more than one element set
Four frameworks matter to most operators, and they overlap heavily without being identical. Knowing which one you are being judged against decides what evidence you have to hold.
The United States regulation, 29 CFR 1910.119, is the oldest and the most widely copied. It defines 14 elements and it is prescriptive about several of them, including the five year revalidation cycle for process hazard analysis and the written requirements for mechanical integrity. Its structure is duty based. Each element is something the employer shall do.
- Employee participation
- Process safety information
- Process hazard analysis
- Operating procedures
- Training
- Contractors
- Pre startup safety review
- Mechanical integrity
- Hot work permit
- Management of change
- Incident investigation
- Emergency planning and response
- Compliance audits
- Trade secrets
The CCPS Risk Based Process Safety framework, published in 2007, reorganises the same ground into 20 elements across 4 pillars. It is not a regulation and nobody audits you against it directly, which is exactly why it is more useful for design. It names the things the regulation assumes and never states, in particular culture, competency and management review, and it grades effort by risk rather than applying one intensity everywhere.
- Commit to process safety, covering culture, compliance with standards, competency, workforce involvement and stakeholder outreach
- Understand hazards and risk, covering process knowledge management and hazard identification and risk analysis
- Manage risk, covering operating procedures, safe work practices, asset integrity and reliability, contractor management, training and performance assurance, management of change, operational readiness, conduct of operations and emergency management
- Learn from experience, covering incident investigation, measurement and metrics, auditing, and management review and continuous improvement
In Europe the Seveso III Directive 2012/18/EU requires a safety management system as part of the major accident prevention policy, with the content set out in its Annex III. The emphasis differs from the American model. Seveso pushes harder on demonstrating that the risk has been assessed and controlled, and it makes land use planning around the establishment an explicit obligation rather than an afterthought.
In India the governing instruments are the Manufacture, Storage and Import of Hazardous Chemical Rules 1989 made under the Environment Protection Act, together with section 41B of the Factories Act 1948 for hazardous processes. Between them they require notification of major accident hazard installations, a safety report, an on site emergency plan, and disclosure of information to workers and to the public living nearby. The obligation to keep the safety report current after modification sits inside those rules and is the part most often missed.
29 CFR 1910.119 and its appendix A, CCPS Guidelines for Risk Based Process Safety 2007, Directive 2012 18 EU annex III, and the MSIHC Rules 1989 with the Factories Act 1948 section 41B. Where a facility falls under more than one, build to the most demanding and map the others onto it rather than maintaining parallel systems.
Does the regime apply to you at all
Applicability is a quantity question and it is answered chemical by chemical, not site by site. Under the American rule the trigger is either a listed highly hazardous chemical held at or above its appendix A threshold quantity, or 10,000 pounds of a flammable liquid or gas in one process. Under Seveso the establishment is lower tier or upper tier depending on the quantities in Annex I, with a summation rule for mixtures of hazards. Under the Indian rules the schedules perform the same function.
Two consequences follow that are worth stating plainly. The first is that a process just below the threshold is not safe, it is unregulated. The chemistry does not know where the line is. The second is that the quantity in one process governs, and process is defined broadly enough that interconnected vessels which could reasonably be involved in a single release are counted together. Splitting an inventory across nominally separate tanks that share a manifold rarely survives scrutiny.
Treating the threshold as the safety boundary is the most common structural error in a young system. Threshold quantities decide the legal duty. The hazard analysis decides what the plant actually needs, and it should be run on any process that can kill someone regardless of what the schedule says.
What each element has to produce
The useful test for an element is not whether a procedure exists. It is what artefact the element produces, who consumes that artefact, and how anyone would know it had gone stale. An element that produces nothing another element uses is decoration.
The knowledge elements produce the description of the plant and the assessment of what it can do. Process safety information produces the chemical data, the technology basis and the equipment basis, including the design codes, the relief design basis and the drawings. Hazard identification and analysis consumes all of that and produces a recorded set of scenarios, safeguards and recommendations with owners and dates. Every quantitative study downstream, from relief sizing to consequence modelling, is calculated from these two.
The operating elements produce the controls that keep the plant inside the envelope. Operating procedures produce the safe operating limits and the consequences of deviation, not just the sequence of steps. Training and competency produce demonstrated capability against those procedures rather than attendance records. Asset integrity produces an inspection and test plan derived from the failure mechanisms that actually apply, together with the results. Safe work practices produce the permits, the isolation certificates and the confined space and hot work controls. Operational readiness, called pre startup safety review in the American rule, produces the record that the plant was verified before hazardous material was introduced.
Management of change sits across both groups and is the element with the widest reach. It produces the technical review of a proposed change, the authorisation, the updates to every artefact the change invalidates, and the confirmation that those updates happened before the change went live. A management of change record that closes while its drawing and procedure updates are still outstanding has produced an authorisation and nothing else.
The learning elements produce the evidence that the system is working. Incident investigation produces causes at a system level and actions that address them. Metrics produce the indicators discussed below. Audits produce findings with owners. Management review produces decisions, resources and a documented judgement about whether the risk is being controlled, which is the element that most systems either skip or reduce to a presentation.
Measuring whether it works
Counting injuries tells you almost nothing about process safety. A site can run for years with an excellent recordable injury rate and a deteriorating containment record, because the two measure different things. ANSI API RP 754 exists to separate them, and it is now the common language for process safety performance across the sector.
The recommended practice arranges indicators in four tiers, drawn as a pyramid. Tier 1 is a loss of primary containment of greater consequence, judged by the material, the quantity released and the outcome. Tier 2 is a loss of primary containment of lesser consequence. Tier 3 covers challenges to the safety system, such as demands on a safety instrumented function, relief device activations and excursions beyond a safe operating limit. Tier 4 covers operating discipline and management system performance, such as procedure currency, training completion and the closure of overdue actions.
Tier 1 and Tier 2 are lagging indicators. Something has already come out. Tier 3 and Tier 4 are leading indicators, and they are the ones a management review should spend its time on, because they move before the containment record does.
Rates are normalised to 200,000 work hours, which represents 100 full time equivalents working 2,000 hours in a year. Employee, contractor and subcontractor hours are all included, because a contractor breaking containment produces the same release as an employee doing it.
- Tier 1 process safety event rate
- count of Tier 1 events in the period
- total employee, contractor and subcontractor work hours in the same period
The recommended practice also allows a severity weighted rate, in which each event carries a score reflecting the seriousness of its outcome rather than counting one for one. This matters at sites where a handful of small events would otherwise read the same as one serious one.
- severity weighted process safety event rate
- severity score assigned to event i
The arithmetic is trivial and the interpretation is not. A site with a small workforce generates few hours, and a single event moves its rate a long way. The calculator below shows both the rate and how far one additional event would move it, because that second number is what tells you whether the first one means anything.
Process safety event rate calculator
Tier 1 and Tier 2 rates on the API RP 754 basis, with the volatility that sits behind them.
One more Tier 1 event in this period would take the rate to 0.75, a rise of 0.25 or 50%. At the current count the site works 400k hours per Tier 1 event. If a single event moves the rate by more than a quarter, a period on period comparison is describing chance rather than performance, and the Tier 3 and Tier 4 indicators will tell you more than this number will.
Rates assume a consistent event definition across the period and complete work hour capture including contractors. Both assumptions fail more often than the arithmetic does. A rate that improves in the same quarter that reporting standards tightened is measuring the reporting, not the plant.
Where systems fail in practice
Systems rarely fail by missing an element. They fail by producing the artefact and then letting it go stale, and the pattern is consistent enough to check for directly.
Action closure is the most measurable failure. Hazard study recommendations, incident actions and audit findings all produce a queue, and the health of that queue is a leading indicator in its own right. The number worth watching is not how many are open but how old the oldest one is, and whether any have been closed by revising the action rather than doing it.
Recurring audit findings are the second. A finding that appears in three consecutive audit cycles is not a finding about the topic. It is a finding about management review, because something has been reported, accepted and not resourced three times.
Competence drift is the third and the hardest to see. Training records show completion. They do not show whether the people who understood why a limit exists have retired, and whether the people who replaced them inherited the limit as a number on a screen with no explanation attached. Elements that depend on judgement, particularly hazard analysis and management of change screening, degrade quietly when this happens.
The fourth is the currency of the process safety information itself. Because every other element reads from it and none of them re verify it, an error there propagates into the hazard study, the relief design basis, the area classification schedule and the isolation list at the same time.
Running a gap assessment
A gap assessment that scores each element out of five and produces a spider diagram is easy to commission and almost impossible to act on. A useful one is evidence based and asks the same three questions of every element.
- What does this element produce, and can you show me the most recent one
- Who consumes it, and can they show me that they used the current version
- How would anyone find out that it had gone stale, and when did that check last run
The second question is the one that finds real gaps, because it tests the joint between elements rather than the element itself. Most systems are stronger inside their elements than between them.
Score against a maturity scale rather than a compliance tick, because compliance is binary and improvement is not. A workable scale runs from absent, to present but informal, to documented and applied, to measured, to reviewed and improved. Most elements at a mature site sit at documented and applied, and the gap that matters is usually the step to measured.
Weight the findings by risk. An informal management of change process on a plant handling a toxic gas under pressure is a different finding from the same informality on a utilities system, and a report that ranks them together has wasted the reader's attention.
What a working system looks like
A process safety management system is working when three things are true at once. The description of the plant matches the plant. Every study that depends on that description states which version it used. And the queue of things the system has told you to fix is getting shorter rather than older.
None of that requires a larger system. It requires that each element produce something another element actually consumes, and that somebody is accountable for the joint between them.