Appendix D. Proof of Theorem 2
Membership in
is immediate. Given an instantiation
of
, it is easy to verify if
is a solution by querying the PP-oracle if
which is a problem known as D-MAR [
15]. To prove hardness, we show that E-MAJSAT [
56] can be reduced to D-Reverse-MAP in polynomial time, based on a slight modification of the reduction to classical MAP proposed in [
34]. The E-MAJSAT problem is defined as follows. Given a Boolean formula
over Boolean variables
: Is there an instantiation
of
such that the majority of instantiations
of
satisfy
(formula
holds at
)? We show that we can answer E-MAJSAT by answering D-Reverse-MAP on an SCM
that simulates the formula
and that can be constructed efficiently. The SCM
is constructed inductively, as shown in [
34] (Ref. [
34] intended to construct a Bayesian network, but their construction is an SCM since all internal nodes in the network have functional CPTs), and always has a single leaf node, denoted
The construction is based on three rules: (1) If
, then
has a single binary node
X with values
and a uniform prior so
; (2) If
, then
is constructed from
by adding a binary node
as a child of
with structural equation
; and (3) If
(
), then
is constructed from
and
by adding a binary node
as a child of
and
with structural equation
(
). We are now ready for the last step of the proof. Given a Boolean formula
over variables
, and given its SCM
that has distribution
, we next show that there is an instantiation
such that
(D-Reverse-MAP query) iff there is an instantiation
such that the majority of instantiations
of
satisfy
(E-MAJSAT query). Let
and
denote
. By construction of
[
34], we have
for all instantiations
;
if
and
otherwise. Then
if
and
otherwise. We finally have:
Now that we have shown membership and hardness, D-Reverse-MAP is
-complete.
Appendix G. Proof of Lemma 1
This proof uses the elimination concepts and notations reviewed in
Appendix F.
Let be the moral graph of and be the cluster induced by eliminating variable from . Let n be the number of variables in G. To prove Lemma 1, it suffices to prove the following statement: for which we prove next by induction.
Let denote the neighbors of X in and let denote the neighbors of X in . Let denote the children of H in . For each , let denote the parents of Z in G. First, we show the statement holds when . When creating from , the introduction of node H would cause two classes of edges that do not exist in to be added to : (1) for ; (2) if Y is a parent of some node , that is Y and H are common parents of some node Z. This means that before the elimination starts, for any node X, if , we have ; otherwise . Hence, . Consider now the elimination of assume that the statement holds for . We observe that if , then the elimination of would cause only one type of additional edges to be added to , that is for . This is because the elimination of will form a clique among in , but already forms a clique after is eliminated from . This implies that eliminating () will never cause any additional edge to be added among any two nodes that are both not H (in other words, all additional edges added are incident on H). Thus, before we eliminate , we have and this implies which concludes the proof.
Appendix I. Proof of Theorem 6
This proof uses the elimination concepts and notations reviewed in
Appendix F.
For a node X in an SCM G, we use to denote its duplicate in an n-world model of G. If node X is shared between all n worlds, then for all k. For a set of variables , we use to denote .
Let G be an SCM, be a subset of its roots (unit variables), and let be a corresponding objective model with n components. Our proof is based on constructing an augmented objective model by adding edges to and then showing that the bounds of Theorem 6 hold for . Our proof is based on Lemmas A2 and A4, which we formally state and prove later:
- –
Lemma A2 complements Theorem 4 by showing that any -constrained elimination order for an SCM can be converted into a -constrained elimination order for a corresponding n-world model while preserving the width of the order.
- –
Lemma A4 concerns the augmentation of an SCM by a root node H and some edges that originate from H. In particular, given a -constrained elimination order of width w for the SCM, the lemma shows how to construct a -constrained elimination order for its augmentation with width .
We start by showing how to construct the augmented objective model
from
. Let
H be the mixture node of
and
be the set of all outcome variables in the objective function of Equation (
2). We obtain
by adding to
an edge
for each
if such an edge does not already exist in
. The edges of
are a superset of the edges of
, so it suffices to show that the bounds of Theorem 6 hold for
. We will next use
to denote a triplet (3-world) model of
G. We will also use
to denote the augmentation of
with mixture node
H and edges
for
. Note that the augmented objective model
corresponds to
n copies of
that share root nodes
. Hence,
is an
n-world model of
.
Let be a -constrained elimination order for G with width w. Since is a triplet (3-world) model of G, Theorem 4 tells us that there exists an elimination order of with width such that (order will also be -constrained). Recall that is obtained from by adding a root node H and some edges that emanate from H. By Lemma A4, there exists an elimination order for with width such that . Moreover, if the objective function has a single outcome variable Y, then H has a single child Y in , so, also by Lemma A4, we have . Since is an n-world model of based on roots of , we have by Lemma A2. In summary, we have . If the objective function has a single outcome variable, we have . This concludes the proof of Theorem 6.
We will next formally state and prove Lemmas A2 and A4, which we used in the above proof.
Lemma A2. Consider an SCM G, a subset of its roots, and a corresponding n-world model for G that shares . If w is the width of a -constrained elimination order π for G, and is the width of the corresponding -constrained elimination order for , then .
Given an elimination order
for an SCM
G, we can convert it into a corresponding elimination order
for its
n-world model
(referenced in the above lemma) by replacing each variable
in
with its duplicates
, as in Definition 2 in [
37]. If
is
-constrained, then
will also be
-constrained. Moreover, we define a graph sequence for the
n-world model
where
is the moral graph of
, and
is obtained by eliminating all duplicates of variable
, i.e.,
, from
.
Proof. Suppose we eliminate variables from using orders . We claim that at every elimination step i, the following properties hold:
- (A)
For each node ,
- (B)
For each node ,
We next show that properties (A), (B) imply and then prove these properties. Let . If , then when its duplicate is eliminated from , we have . If , then when Y is eliminated from , we have since all non-shared nodes have been eliminated before Y, i.e., . This means that the cluster induced by eliminating a variable from always has the same size as the cluster induced by eliminating the corresponding variable from G, which implies .
We next prove properties (A), (B) by induction. By definition of an n-world model, these properties hold initially for . Suppose they hold for and consider . Let . Then is the result of eliminating nodes from . We consider two cases.
Case: . Consider each node Z in . If Z is not a neighbor of in , then will not be affected by the elimination of and the properties hold by the induction hypothesis. Otherwise, node Z falls into two cases: (A) a duplicate of a node , (B) a shared node .
- (A)
by the induction hypothesis, neighbors of
in
must belong to the
k-th world, so
can only be a neighbor of the
k-th duplicate
. This means that
can only be affected by the elimination of
. By the definition of Variable Elimination, we have:
This proves property (A).
- (B)
by the induction hypothesis,
U must be a neighbor of all duplicates
. We have:
This proves property (B).
Case: . In this case, only contains nodes in . Property (A) holds trivially. And the relation in property (B) reduces to . By the induction hypothesis, we know and thus . Property (B) holds. This concludes the proof. □
The proof of Lemma A4 requires the following result on eliminating variables from graphs.
Lemma A3. Consider a DAG G, a subset of its nodes, and a node H in G where . Let be the moral graph of G, and be the result of eliminating all nodes other than from . For any node , X is adjacent to H in if and only if there exists a path between X and H in that does not include a node in .
Proof. We first prove the if direction. Suppose there exists such a path in . Eliminating node Y from will lead to a path . Since nodes in cannot appear along this path, eliminating all nodes other than will lead to the edge in . We next prove the only if direction by contraposition. Suppose there is no path between X and H in that does not include a node in . There are two cases: (1) there is no path between X and H; (2) every path between X and H includes at least one node , which has the form . In the first case, X and H will be disconnected in . In the second case, eliminating all nodes other than from such paths will lead to , so X cannot be directly adjacent to H in . This concludes the proof. □
Lemma A4. Consider an SCM G and a subset of its roots. Suppose SCM is obtained from G by adding a root node H as a parent of some nodes in G, where . Let π be a -constrained elimination order for G, and let be a -constrained elimination order of obtained from π by placing H just before variables . If π has width w and has width , then . Moreover, if H has a single child in , then .
Proof. Let denote variables other than in G, and let . Suppose we first eliminate variables , then H, and finally from using order . This results in a graph sequence where and . Here, is obtained by eliminating all variables from , and is obtained by eliminating H from . We claim:
if , then for each node in , we have .
if , then for each node in , we have . Moreover, if H has a single child in , then .
We first show that the above claim implies the lemma, and then follow by proving the claim. Suppose we are eliminating variable Y from . If , then and the above claim implies . If then and the above claim implies , and when H has a single child. This guarantees the statement of the lemma: , and if H has a single child in .
We next prove our claim by induction. Let denote the children of H in . When constructing the moral graph from , the introduction of node H causes two classes of edges that do not exist in to be added to : for , and if Y is a parent of some node , that is Y and H are common parents of some node Z. All of these extra edges are incident on H, meaning that for any node X in , . Thus, our claim holds for . Next, assume our claim holds for (induction hypothesis) and consider . We have two cases.
Case: . Let
. Consider each node
X in
. If node
X is not a neighbor of
Y in
, then
X is not affected by the elimination of
Y, i.e.,
and
. So the claim holds by the induction hypothesis. Otherwise, we can bound
as follows:
Case: . For this case, only contains nodes in . It is trivial that for each node X in . Recall that eliminating H from results in . By the induction hypothesis, all extra edges in that do not exist in must be incident on H. Consider the special case where H has a single child in . We claim that in this case, every two nodes in are adjacent in , meaning that the neighbors of H already form a clique in . Thus, eliminating H from will not add any fill-in edges in . This guarantees , i.e, for all . We finally turn to proving this claim by contradiction. Suppose that node and are neighbors of H in but are not adjacent in . By Lemma A3, in , there must be a path between and H that does not include nodes in , and a path between that does not include nodes in . Since H is a root and only has one child Z in , must have the form in and must have the form in . Thus, there must be a path in that does not contain nodes in . By Lemma A3, after eliminating all nodes other than from , and must be adjacent in . This leads to a contradiction. □