Skip to contents

Ce guide existe aussi en français : Introduction à ggfacto.

ggfacto draws the geometric data analyses of FactoMineR so that they can be read without losing sight of the data. Hover over a point of the interactive graph, and it shows the crosstables it comes from, each percentage coloured by its deviation from the mean: blue where a category is over-represented, red where it is under-represented. A category at the edge of the cloud then reads as the sum of its deviations, which is the best guard against over-interpreting the geometry.

The three methods share one workflow:

  1. the analysis, with the data frame first, as in tabxplor::tab();
  2. interpret(), the table of the axes, with the eigenvalues underneath;
  3. ggfacto(), the graph, made interactive by interactive = TRUE;
  4. the clusters, added to the data frame with mutate() and described by clust_tab().

One option decides how every table prints, tabxplor’s and ggfacto’s alike: with options(tabxplor.print = "html"), they print as html tables in RStudio, Positron or a document.

Principal component analysis: means

A PCA summarises numeric variables. Here, the cars of mtcars are described by six active measurements, and the number of cylinders is a supplementary variable.

cars <- mtcars |> mutate(cyl = factor(cyl))

pca <- principal_component_analysis(cars, c(mpg, disp, hp, drat, wt, qsec))
interpret(pca)
Variables Axe 1 Axe 2 Axe 3
variable mean sd sd/mean coord contrib cos2 coord contrib cos2 coord contrib cos2
<mean> <sd> <cv> <mean> <col%> <row%> <mean> <col%> <row%> <mean> <col%> <row%>
mpg 20.09 5.93 30% -0.94 21% 88% -0.06 0% 0% -0.11 4% 1%
disp 230.72 121.99 53% 0.95 22% 91% 0.06 0% 0% 0.06 1% 0%
hp 146.69 67.48 46% 0.87 18% 76% -0.39 13% 15% 0.08 2% 1%
drat 3.60 0.53 15% -0.75 13% 56% -0.47 19% 22% 0.46 64% 21%
wt 3.22 0.96 30% 0.90 19% 81% 0.32 9% 10% 0.24 17% 6%
qsec 17.85 1.76 10% -0.52 6% 27% 0.82 58% 67% 0.20 12% 4%
Total 100% 100% 100%
Coordinate on the axis, i.e. its correlation with it: a variable above, by +0.1; +0.2; +0.4; +0.8 SD; a variable below, by -0.1; -0.2; -0.4; -0.8 SD.
contrib: its contribution to the variance of the axis; an axis sums to 100 %
cos2: quality of representation
sd/mean: coefficient of variation – the standard deviation as a percentage of the mean, comparable between variables measured in different units
Variance
Axe eigenvalue % variance cumul.
<var> <col%>
Axe 1 4.187 69.8% 69.8%
Axe 2 1.148 19.1% 88.9%
Axe 3 0.333 5.6% 94.5%
Axe 4 0.154 2.6% 97.1%
Axe 5 0.125 2.1% 99.1%
Axe 6 0.052 0.9% 100%
Total 6.000 100%

The table opens on the variables as they were before the analysis: mean, standard deviation and coefficient of variation (sd/mean, the standard deviation as a percentage of the mean, which can be compared across variables measured in different units). Then comes each axis: the coordinate of each variable (its correlation with the axis, coloured), its contribution and its quality of representation (cos2). The first axis, 70 % of the variance, sets displacement, horsepower and weight against fuel economy (mpg): a size effect. The second singles out the cars that are slow off the mark (qsec).

ggfacto(pca, cars, sup_vars = cyl, interactive = TRUE)

The biplot draws the cars (in grey), the active variables (the arrows), and the supplementary categories at the centre of gravity of their cars. Hover: an arrow gives its variable’s mean and coefficient of variation; a supplementary category gives the mean of every active variable within it, coloured by its deviation from the overall mean; a car gives its own values. ggfacto(pca, profiles = FALSE) draws the correlation circle alone.

Correspondence analysis: a crosstable

A CA is the analysis of one crosstable, so it is computed on the table itself. Here, religion is crossed with party identification in the US General Social Survey sample of forcats::gss_cat, non-responses left out.

gss <- forcats::gss_cat |>
  filter(!relig %in% c("No answer", "Don't know", "Not applicable"),
         !partyid %in% c("No answer", "Don't know"))

tab(gss, relig, partyid, pct = "row", color = "contrib")
partyid
relig Other party Strong
republican
Not str
republican
Ind,near rep Independent Ind,near dem Not str
democrat
Strong democrat Total
<row%> <row% (n)>
Inter-nondenominational 1% 13% 9% 8% 26% 12% 10% 20% 100% (   108)
Native american 0% 0% 4% 4% 26% 17% 30% 17% 100% (    23)
Christian 3% 8% 15% 10% 22% 9% 18% 14% 100% (   684)
Orthodox-christian 0% 10% 12% 11% 18% 15% 20% 15% 100% (    94)
Moslem/islam 1% 2% 3% 7% 21% 19% 28% 18% 100% (    99)
Other eastern 0% 6% 6% 9% 38% 22% 9% 9% 100% (    32)
Hinduism 0% 0% 9% 3% 38% 17% 26% 7% 100% (    69)
Buddhism 4% 6% 3% 5% 21% 27% 20% 13% 100% (   143)
Other 7% 3% 7% 12% 22% 21% 14% 15% 100% (   221)
None 3% 3% 8% 8% 28% 18% 16% 16% 100% ( 3 509)
Jewish 2% 6% 10% 4% 10% 14% 21% 32% 100% (   388)
Catholic 1% 9% 14% 9% 21% 11% 21% 14% 100% ( 5 090)
Protestant 1% 15% 17% 8% 16% 10% 16% 17% 100% (10 794)
Total 2% 11% 14% 8% 19% 12% 17% 16% 100% (21 254)
pvalue (Chi2 !) <0.01%
Cramér’s V 0.10
Contribution to Chi2: cell over-represented vs independence, by ×1; ×2; ×5; ×10 the mean contribution; cell under-represented, by ×1; ×2; ×5; ×10 the mean contribution.

With color = "contrib", a cell is coloured by its contribution to the table’s variance, which is what weighs in the analysis: blue when it is over-represented, red when under-represented.

ca <- correspondence_analysis(tab(gss, relig, partyid))
interpret(ca)
Variable Positive_levels Negative_levels
<col%> <col%>
Axe 1: 76.2%
of variance
relig None 55.0% Protestant 32.4%
relig: above mean ctr 55.0% 32.4%
partyid Independent 21.1% Strong republican 35.9%
Ind,near dem 18.3% Not str republican 18.0%
partyid: above mean ctr 39.4% 53.9%
Axe 2: 10.9%
of variance
relig Jewish 43.0% Catholic 34.3%
relig: above mean ctr 43.0% 34.3%
partyid Strong democrat 46.4% Independent 17.2%
Other party 14.3%
partyid: above mean ctr 60.8% 17.2%
Contribution to the variance of the axis: a level on the positive side, contributing ×1; ×2; ×5; ×10 the mean contribution; a level on the negative side, contributing ×1; ×2; ×5; ×10 the mean contribution.
Axe eigenvalue % variance cumul.
<var> <col%>
Axe 1 0.049 76.2% 76.2%
Axe 2 0.007 10.9% 87.1%
Axe 3 0.005 7.5% 94.6%
Axe 4 0.002 2.7% 97.3%
Axe 5 0.001 2.0% 99.4%
Axe 6 0 0.5% 99.9%
Axe 7 0 0.1% 100%
Total 0.064 100%

On each axis, interpret() keeps only the categories that contribute more than average, facing each other by the sign of their coordinate. The first axis, 76 % of the variance, opposes Republican Protestants to the non-religious, who lean towards independents and Democrats; the second sets Jewish respondents and strong Democrats apart from Catholics. Once interpreted, the axes are named, and the names are printed on the graphs and in the tables.

ca <- name_axes(ca, "Republican Protestants / no religion",
                    "Catholics / Jewish Democrats")
ggfacto(ca, interactive = TRUE)

Hover over a category to see its profile: its row of the crosstable, each percentage coloured by its deviation from the mean. The CA shows the structure of the table’s deviations, and says nothing of their size: that is what the coloured table beside it is for.

Multiple correspondence analysis: response patterns and the Burt table

An MCA analyses several questions at once. The tea survey of FactoMineR asked 300 people 18 active questions about how they drink tea.

data(tea, package = "FactoMineR")

mca <- multiple_correspondence_analysis(tea, 1:18)
interpret(mca, axes = 1:2)
Question contrib Positive_levels Negative_levels
<col%> <col%> <col%>
Axe 1: 9.9% of variance (mod. 55%) where 15.7% chain store+tea shop 11.3% chain store 4.4%
tearoom 13.9% tearoom 11.2% Not.tearoom 2.7%
how 11.2% tea bag+unpackaged 6.8% tea bag 4.3%
friends 9.1% friends 3.2% Not.friends 6.0%
resto 8.5% resto 6.3% Not.resto 2.2%
price 8.1% p_variable 3.5% p_branded 3.0%
tea.time 7.2% tea time 3.1% Not.tea time 4.1%
pub 5.5% pub 4.4%
work 4.2% work 3.0%
How 3.9% other 2.3%
Tea 3.4% green 3.0%
lunch 2.8% lunch 2.4%
Above mean ctr 87.0% 57.3% 29.6%
Axe 2: 8.1%
of variance
(mod. 28%)
where 28.6% tea shop 23.9% chain store 4.6%
price 25.6% p_upscale 20.5% p_branded 2.5%
how 23.4% unpackaged 18.9% tea bag 4.5%
Tea 7.3% green 3.3% Earl Grey 2.4%
Above mean ctr 80.5% 66.6% 13.9%
Contribution to the variance of the axis: a level on the positive side, contributing ×1; ×2; ×5; ×10 the mean contribution; a level on the negative side, contributing ×1; ×2; ×5; ×10 the mean contribution.
contrib: the whole question’s contribution to the axis
Axe eigenvalue % variance cumul. Benzecri’s
modified rate
cumul. mod.
<var> <col%> <col%>
Axe 1 0.148 9.9% 9.9% 55.4% 55.4%
Axe 2 0.122 8.1% 18.0% 28.1% 83.4%
Axe 3 0.090 6.0% 24.0% 7.6% 91.1%
Axe 4 0.078 5.2% 29.2% 3.3% 94.3%
Axe 5 0.074 4.9% 34.1% 2.1% 96.5%
Axe 6 0.071 4.8% 38.9% 1.6% 98.1%
Axe 7 0.068 4.5% 43.4% 1.0% 99.1%
Axe 8 0.065 4.4% 47.7% 0.6% 99.7%
… of 27
Total 1.500 100% 100%

In an MCA the raw rates of variance are low by construction, so the number of axes to interpret is chosen on Benzécri’s modified rates, in the eigenvalue table: 83 % for the first two axes here. The first axis opposes tea drunk only at home to tea drunk out, with friends — in tearooms, restaurants and pubs; the second, supermarket tea bags to loose tea from specialist shops.

mca <- name_axes(mca, "at home / out, with friends",
                      "supermarket / tea shop")
ggfacto(mca, tea, sup_vars = c(sex, SPC), interactive = TRUE)

The data frame is passed again, second, for the supplementary variables (in italics). Hover, and two things can be read:

  • the grey points are the response patterns: the people who gave exactly the same answers, sized by their number. Each one lists its answers;
  • each category shows its crosstables with every other active question. Together they form the Burt table the MCA is computed from, its deviations from the mean coloured. The data is back in view, without crossing the 18 questions two by two.

Clustering

Hierarchical clustering groups together the individuals closest to one another on the first axes. Without nb_clust, it draws the tree and cuts it where the gain in between-cluster variance drops.

hierarchical_clust(mca, ncp = 3)

The clusters are added to the data frame with mutate(), named in the order of your choice ("name" = number): the tree is not built again.

tea <- tea |>
  mutate(clusters = hierarchical_clust(mca, ncp = 3, names = c(
    "Supermarket tea bags"   = 1,
    "Sweet Earl Grey"        = 2,
    "Earl Grey with friends" = 4,
    "Black, no sugar"        = 5,
    "Tea-shop connoisseurs"  = 3,
    "Tea everywhere"         = 6
  )))

clust_tab(mca, tea, clusters)
clust
lvs Supermarket
tea bags
Sweet Earl Grey Earl Grey with
friends
Black, no sugar Tea-shop
connoisseurs
Tea everywhere Ensemble
<col%> <col%>
breakfast 65% 12% 44% 61% 44% 67% 48%
Not.tea time 49% 79% 34% 15% 50% 14% 44%
evening 10% 50% 41% 30% 31% 58% 34%
lunch 6% 9% 23% 6% 9% 39% 15%
dinner 4% 11% 6% 3% 22% 0% 7%
always 13% 39% 39% 27% 44% 64% 34%
home 100% 86% 100% 100% 97% 100% 97%
Not.work 81% 88% 59% 73% 78% 36% 71%
Not.tearoom 97% 98% 89% 45% 75% 39% 81%
friends 30% 68% 86% 76% 56% 100% 65%
Not.resto 90% 80% 66% 67% 91% 33% 74%
Not.pub 97% 84% 70% 82% 78% 44% 79%
black 39% 2% 6% 73% 25% 17% 25%
Earl Grey 44% 84% 92% 24% 47% 81% 64%
green 16% 14% 2% 3% 28% 3% 11%
alone 61% 88% 67% 52% 69% 44% 65%
lemon 4% 9% 9% 6% 19% 31% 11%
milk 35% 4% 20% 24% 12% 22% 21%
other 0% 0% 3% 18% 0% 3% 3%
No.sugar 63% 18% 50% 85% 56% 47% 52%
tea bag 84% 82% 64% 24% 9% 17% 57%
tea bag+unpackaged 15% 12% 31% 61% 19% 81% 31%
unpackaged 1% 5% 5% 15% 72% 3% 12%
chain store 95% 96% 67% 33% 6% 19% 64%
chain store+tea shop 5% 0% 31% 67% 9% 81% 26%
tea shop 0% 4% 2% 0% 84% 0% 10%
p_branded 49% 50% 27% 9% 6% 17% 32%
p_cheap 4% 5% 2% 0% 0% 0% 2%
p_private label 15% 11% 2% 3% 0% 3% 7%
p_unknown 5% 7% 5% 0% 0% 3% 4%
p_upscale 5% 5% 3% 45% 81% 8% 18%
p_variable 22% 21% 62% 42% 12% 69% 37%
% of population 26% 19% 21% 11% 11% 12% 100%
n 79 56 64 33 32 36 300
Percentage points (risk) difference: cell ≥ the Total column +5; +10; +20; +30 points; cell ≤ the Total column -5; -10; -20; -30 points.

The table describes each cluster by the active variables, weighted as in the analysis, with the deviations from the whole population coloured. The graph places the clusters in the cloud, with their individuals in their colour:

ggfacto(mca, tea, clust = clusters)

Going further

  • Survey weights: wt = in all three analyses. The weights carry through to the crosstables in the tooltips.
  • Specific MCA: excl = leaves categories out of the analysis (missing values, by default). Subpopulation: filter =, or filter() in the pipe, then the whole data frame passed to ggfacto().
  • Graphs: they are ggplot objects, which you can extend with + before ggi(). ggsave2() saves them, and ggmca_3d() and ggpca_3d() draw them in three dimensions.
  • The coordinates of the individuals are added to the data frame with mutate(axis1 = axis_coord(mca, 1)), just like the clusters.

The reference documents every function.