The write-up is here. As always, the motivation is below (it’s quite lengthy this time; maybe the longest on this blog, so far).
Consider a mean-zero function \(f\) on a convex, bounded open set \(\Omega\) (or, more generally, a domain \(\Omega\) containing a ball \(B\) such that for every \(x \in B\), the set \(\Omega\) is star-shaped about \(x\)). When is it the case that $$\text{div}(u) = f$$ in \(\Omega\), for some function \(u\) vanishing on the boundary \(\partial \Omega\)? (The mean-zero condition, of course, is for consistency, recalling the divergence theorem.)
As it turns out, it is a construction of Bogovskiĭ furnishes an explicit example of a right-inverse to the divergence operator: on the linear space \((L^p(\Omega))_0 = \{g \in L^p(\Omega) : \int_{\Omega} g \, d\lambda^d = 0\}\), \(1 < p < \infty\), there exists a linear operator mapping scalar functions to vector fields, \(T : (L^p(\Omega))_0 \to [W^{1, p}_0(\Omega)]^d\), such that $$\text{div}(T(f)) = f$$ for each \(f \in (L^p(\Omega))_0\). The fact that such an operator even exists, even for smooth functions, is very much not obvious, and somewhat surprising (at least to me).
The process starts with an integral representation formula, which can be shown to map functions in \(C^{\infty}_c(\Omega)\) of mean-zero to \([C^{\infty}_c(\Omega)]^d\) and satisfy the relevant equation distributionally. Furthermore, the operator we get has weakly-singular kernel of homogeneity \(– (d \ – \, 1)\), so by HLS it maps \((L^p)_0\) to \(L^{p^*}\), which here denotes whatever improved space is produced by fractional integration (a.k.a. the Riesz potential).
After a lot of irritating computation (involving some very messy integrations-by-parts due to the singularity; this is reminiscent of a species of calculation common in classical treatments of the Newtonian potential), one can also get a pointwise representation for the partial derivatives of \(T(f)\), which is a (true) singular integral.1
So to get \(W^{1, p}\) boundedness of the right-inverse \(T\), one need only prove \(L^p\) bounds on the derivatives (which are given by singular integral operators acting on the input \(f\)).
Needless to say, I find this an interesting problem: the possibility of controlling the output of the divergence operator is an intriguing question theoretically (and useful in the context of fluids, some friends tell me), and having quantitative norm bounds for this smoothing operator is quite nice. This also fits in to a larger web of results giving equivalences of spaces of functions whose distributional divergence, gradient, or curl are known to enjoy various specific properties. One can search for the work of C. Amrouche, with P. G. Ciarlet and C. Mardare.2
On the question of \(L^p\) boundedness for these singular integrals, Calderón-Zygmund theory provides the answer. The relevant singular integrals will be of a particular form, of non-convolution type and yet also falling outside the modern framework. To bound these operators, we will require a specific result: one of the major theorems of Calderón and Zygmund’s second paper.3 (Recall that their first paper, in Acta, was a historical milestone that illustrated the power of real-variable methods in singular integral theory, and pioneered the analysis of singular integrals of convolution type.)
Theorem: Let \(N : \mathbb{R}^d \times (\mathbb{R}^d \setminus \{0\}) \to \mathbb{C}\) is a (product) measurable function, positively-homogeneous of degree \(– d\) in the second argument. Suppose the following hold:
- For every \(x \in \mathbb{R}^d\), \(y \mapsto N(x, y)\) is absolutely-integrable on \(S^{d \ – \, 1}\), with mean-value zero; and
- There exists \(1 < q < \infty\) such that \(\sup_{x \in \mathbb{R}^d} \|N(x, \cdot)\|_{L^q(S^{d \ – \, 1})} < \infty\).
Then for the kernel \(K : (\mathbb{R}^d \times \mathbb{R}^d)^* \to \mathbb{C}\) defined by $$K(x, y) = N(x, x \ – \, y),$$ the truncated singular integral operators $$(T_{\epsilon} f)(x) = \int_{|x \ – \, y| > \epsilon} K(x, y) f(y) \, dy$$ obey $$T_* f = \sup_{\epsilon > 0} |T_{\epsilon} f| \in L^p(\mathbb{R}^d)$$ whenever \(f \in L^p(\mathbb{R}^d)\) for a \(q / (q \ – \, 1) \leq p < \infty\), and with $$\|T_* f\|_{L^p(\mathbb{R}^d)} \lesssim_{d, p, q, N} \|f\|_{L^p(\mathbb{R}^d)}.$$ Moreover, as \(\epsilon \to 0\), the truncated operators \(T_{\epsilon} f\) converge pointwise a.e. and in \(L^p(\mathbb{R}^d)\) to a limit \(T f\).
This is one of the characteristic results of Calderón and Zygmund’s program; we note that \(T_{\epsilon}\) is not a convolution operator, thus going beyond the first-generation theory, while it is too rough to be a standard kernel (the \(x\) dependence is only measurable), which prevents application of the third-generation theory (with its typical Hölder condition as part of the definition of a standard kernel).
To me, at least, this is a surprising result: again, the dependence on \(x\) is potentially very rough and discontinuous, ruling out the application of smooth-kernel theory, nor is there even any Hörmander-type condition being enforced on \(K\), like in the convolution-operator setting.
The proof of boundedness rests on the method of rotations, which was first introduced in this paper for this purpose. The idea is as follows: for homogeneous kernels, which are fully determined by their behavior on a restriction to any sphere about the origin, one can rewrite the singular integrals in question as modified directional Hilbert transforms, averaged over the unit sphere (over all directions, as one might say). This then allows us to use the one-dimensional theory to rapidly develop a higher-dimensional theory, for many species of singular integral.
The method of rotations also allows us to prove the following result:
Theorem: Let \(N : \mathbb{R}^d \setminus \{0\} \to \mathbb{C}\) be measurable and positively-homogeneous of degree \(– d\); assume \(N \in L^1(S^{d \ – \, 1})\) and the even part \(N_{\text{e}} = \frac{1}{2} (N + \tilde{N})\) of the kernel \(N\) lies in \(L \log L\).
Then for the kernel \(K : (\mathbb{R}^d \times \mathbb{R}^d)^* \to \mathbb{C}\) given by \(K(x, y) = N(x \ – \, y)\), which induces associated truncated and maximal singular integral operators \((T_{\epsilon})_{\epsilon > 0}\), \(T_*\) as above, the same conclusions as the previous result hold for all \(f \in L^p(\mathbb{R}^d)\) for any \(1 < p < \infty\).
This can be generally understood to be the definitive statement for convolution singular integral operators. It is known that for even kernels that are rougher than \(L \log L\), pointwise convergence of the truncations can fail utterly, the maximal operator \(T_*\) fails to be bounded, and the question of boundedness of \(T = \lim_{\epsilon \to 0} T_{\epsilon}\) becomes quite delicate.4
Thus I found it useful to read this paper of Calderón and Zygmund, despite its age. First, it allowed me to receive an introduction to the method of rotations, which I did not encounter in my first course on harmonic analysis. Next, it presents some material that falls outside the scope of the first- and third-generation theory, but is actually quite important in some contexts, as noted above with the construction of the Bogovskiĭ operator. And finally, it develops a piece of theory which, to the best of my knowledge, does not generally receive much coverage from most lecture notes and sources on the subject; the original source, as it turns out, may be the best (and perhaps only) font for this material.5
My write-up of the paper attempts to follow the arguments of the paper fairly closely, but there are some portions which I had to modify extensively. In an effort to make this guide fully self-contained, I included arguments for all the claimed convergence and boundedness, even when this was relegated to other references by the paper. I also tried to adopt a modern, and unified, style of argument in the estimates.
Consequently, this involved excavating some old articles and other texts.6 Below, we’ll describe the structure of the argument, as well as some of the key difficulties, issues, and obstacles in this process of “translation,” in lifting the text from its past style.
- In general, the style of most analysis papers is to focus on the key estimates; measurability is a side issue. I largely agree with this philosophy, but I personally do often want to verify measurability, which can sometimes be unclear for certain objects (e.g., some types of maximal operators applied to general, rough \(L^p\) functions). Some auxiliary constructions here also perform complicated limiting operations on one variable of a function of several arguments, and it’s generally good to verify that (product) measurability is preserved. As is my style, I verify throughout that all the constructions we encounter are measurable.
- There are six major theorems in the paper, which I have enumerated as I through VI. Theorems III and IV are the easiest to show, and most clearly involve the method of rotations; everything becomes an average of a directional Hilbert transform, essentially. Theorems V and VI are more auxiliary results, mostly bounding certain maximal functions with non-radial kernels (so the radial majorant lemma does not apply), but are otherwise well-behaved (and also follow from rotation-like arguments, drawing from the one-dimensional case). Theorems III and IV handle the cases of the two theorems above when \(N\) is odd in the second (singular homogeneous) variable.
- Controlling the even case in the setting of the second theorem (in the convolution setting) is done with an ingenious trick: Calderón and Zygmund use the identity \(\sum_{i = 1}^{d} R_i^2 f = f\), and apply to this the operator \(T\) (or rather, a suitably-modified, smoothly-truncated version thereof). As it turns out, this leads to a “Riesz transform” of the truncated kernel \(N \phi_{\epsilon}\), and one can show that this is well-behaved, in a sense: \(N_2 = \vec{R}(N \phi_{\epsilon})\) is equal to an odd kernel (given by \(N_1 = \vec{R}(N)\) itself; again, this has to be suitably-defined) away from the origin, plus an error that is sufficiently-rapidly-decaying. Near the origin, \(N_2\) can be controlled by a homogeneous function of mean zero, which poses no real problem. The point is that \(N_1\) is now an odd kernel, because \(N\) was even; so Theorems III/IV can be applied to handle this.
- Proving that the modified kernels \(N_1\), \(N_2\) actually existed (that the Riesz transform could be suitably defined on these objects) was irritating. Because the assumption of \(L \log L\) regularity is rather weak, weaker than \(L^q\) for any \(q > 1\), obtaining convergence of the various limits in the definitions of our singular integrals is a real chore. As it turns out, the \(q > 1\) modified case (as in the first theorem, which had the \(L^q\) condition on the sphere) is easier, because one can appeal to certain results regarding the \(L^p\) boundedness of the maximal Riesz transforms for all \(1 < p < \infty\), and convergence quickly follows.
We had to develop the \(L \log L\) theory for the Riesz transform. This took a while, and required going back even further, to read Calderón and Zygmund’s first paper.7 This proceeded by several steps:
First, the development of the \(L \log L\) theory in the first paper is based on certain distributional inequalities. Unfortunately, the distributional inequalities in that paper are stated using the one-dimensional decreasing rearrangement for a function. While this is perfectly valid, the decreasing rearrangement is also somewhat antiquated and (in my view) not fun to work with. This was how some exceptional-set-estimates were stated in the article; I sought to avoid this in my write-up.
Recall that in first introductions to the theory of convolution singular integrals, one typically has the weak-type estimate for singular integrals, stated as follows: $$\lambda^d(\{|T f| > t\}) \lesssim_{d, T} \frac{1}{t^2} \int_{\mathbb{R}^d} |g|^2 \, d\lambda^d + \sum_n \lambda^d(3 Q_n) + \frac{1}{t} \int_{(3 Q_n)^c} |T b_n| \, d\lambda^d,$$ where \(g\) and the \(b_n\) are the good and bad functions obtained from the Calderón-Zygmund decomposition at height \(t\), and the \(Q_n\) are the associated cubes of the bad part. The good-part can then be handled easily, and the contributions from the bad parts are integrable; it is then the case that we get something like $$\lesssim_{d, T} \frac{1}{t^2} \int_{|f| \lesssim t} |f|^2 \, d\lambda^d + \lambda^d \bigg( \bigcup_n Q_n \bigg).$$ Many authors will then use the estimate $$\lambda^d \bigg( \bigcup_n Q_n \bigg) \simeq_d \frac{1}{t} \int_{\bigcup_n Q_n} |f| \, d\lambda^d \leq \frac{1}{t} \int_{\mathbb{R}^d} |f| \, d\lambda^d$$ to conclude the weak-type estimate.
But while this is adequate to get a weak-type \((1, 1)\) bound, this is not an inequality we can integrate in \(t \gtrsim 1\) (which would be what we’d need, if we wanted an \(L^1_{\text{loc}}\) estimate). The difficulty is in the second measure-bound term; \(\frac{1}{t}\) has a logarithmic divergence at infinity. Consequently, we would want something like \(t \leq |f|\) to be in play, to stop the divergence; yet because of the stopping-time structure of the Calderón-Zygmund decomposition, it is not possible to obtain this general lower bound on \(|f|\) on the cubes of the bad part; there is no reason why \(f\) remain sizable throughout the cube.
(Indeed, the whole point of the decomposition is that we stop whenever we have a cube whose associated average goes from \(< t\), for its predecessor, to \(\geq t\) — without investigating further into the fine-scale structure of the function afterwards; we are “zooming in” no further, so to speak. The whole point is that we do not probe too deeply where \(f\) is uncontrollably, generically large and singular. We leave these cubes \((Q_n)\) alone, and they are primarily useful for the size estimate on their union, rather than any pointwise behavior of \(f\) within. The cubes have to be large, in a sense, in order to maintain the condition \(\lambda^d \left( \bigcup_n Q_n \right) t \sim \int_{\bigcup_n Q_n} |f| \, d\lambda^d\); zooming in too close would make the right-hand side far larger than the left.)
As it turns out, retooling the Calderón-Zygmund decomposition is tricky (it’s a classic for a reason). Instead, I separately restricted \(f\) to only the set where it is \(\gtrsim t\) in magnitude, before applying the Calderón-Zygmund decomposition; that is, we only use it to split the high part. The low part is handled by \(L^2\)-boundedness and Chebyshev, as usual (like with the good part of the high part).
Thus my argument turns out to give a nice bound $$\lambda^d(\{|T f| > t\}) \lesssim_{d, T} \frac{1}{t^2} \int_{|f| \leq t} |f|^2 \, d\lambda^d + \frac{1}{t} \int_{|f| \geq t / 8} |f| \, d\lambda^d,$$ and we can integrate this in \(t \gtrsim 1\). Indeed, from this, we precisely get a bound of an \(L^1\) term and a \(L \log L\) term.
With this refined estimate, we can avoid the work with decreasing rearrangements that was so prevalent in the 1952 paper. I really wanted to avoid this because most modern presentations of the theory use a clean form of the Calderón-Zygmund decomposition, and state estimates in terms of integrals over dyadic collections of the underlying function \(f\), without ever mentioning \(f^*\). I really wanted to produce an argument that was in-line with this point of view, and I strived to make it as unified as possible.
Next, to proceed to the “\(L \log L\) implies local integrability” conclusion, we had to show that control in \(L^1\) and \(L \log L\) ensured the limits of the truncated Riesz transforms converged in the \(L^1_{\text{loc}}\) sense. This required a density argument, showing that \(C^{\infty}_c\) functions were dense in the intersection of the two spaces. My proof of density proceeded from direct estimates, using the simple fact that \(x \log(1 + x) \lesssim x^2\) to convert the (annoying) \(|f| \log(1 + |f|)\) expressions into \(|f|^2\) expressions, when appropriate.
But after this, we were able to prove convergence in local \(L^1\); this established existence of the Riesz transforms of the kernels. In preparation for the Theorem II case, with the \(L^q\) estimates, we also estimate the transformed kernels in higher Lebesgue spaces when \(N\) is more regular. This is actually easier, because any \(L^q\) function is significantly more regular than \(L \log L\); furthermore, in this instance, we have pointwise domination by the maximal Riesz transform, and \(L^p\) estimates, which simplify the proof of existence of the limit significantly.
This was a long digression, but it shows some of the steps that were necessary to fill in the gaps in the second paper and make it self-contained, which is something I deeply strove for.
- Furthermore, the preparation for the results of Theorem II require \(L^q\) bounds on the error term that was homogeneous of degree zero. The original \(L^1\) estimates were all right, but the higher \(L^q\) bounds required more. Fortunately, the weakly-singular kernel structure means this is like a Riesz potential, and so we can apply HLS to control this fractional integral. (I now think that the “theorem of Young” that Calderón and Zygmund reference is probably some form of generalized Young’s inequality (the one that allows for weak-\(L^q\) input), or perhaps some form of it, similar to Hedberg’s proof, adapted for fractional integrals. Of course, I have no way of checking this, as 1935 copies of Zygmund’s book are practically impossible to find, even in my university library; and I can find nothing apposite in later editions.)
- Finally, for the proof of Theorem II, it turns out that one can complete the proof simply by treating \(N(x, y)\) as an \(x\)-indexed family of singular kernels: \(N(x, y) = N_x(y)\). Then we apply the same machinery used to prove Theorem I, with the transformations into odd kernels via Riesz transform. Here, due to the separate treatment of the variables, it becomes important to verify measurability; some care is needed to produce a product-measurable limit. (There is quite a bit of trickery; we use Tonelli to interchange the iterated integrals, and one order definitely leaves it unclear why the relevant integral should vanish, but the other order makes it apparent. I’m rather discomfited by this step, but at the end of the day, it does work.) Other than this, however, the calculations do go through, which is nice, and one again returns to the setting of Theorems III and IV; this is how we obtain Theorem II.
Intriguingly, except for the \(L^q\) requirement (for the nonconvolution case) or \(L \log L\) case (for the convolution setting), there isn’t really a strong demand of regularity on the kernels in these theorems; even the size condition \(|K(x, y)| \lesssim |x \ – \, y|^{- d}\), much less the standard-kernel continuity estimates, is not essential here.
This paper was an interesting experience to write. The \(L \log L\) case had me going back to the details for the weak-type \((1, 1)\) bound for Calderón-Zygmund operators; it took an embarrassingly-long amount of experimenting on a Tuesday afternoon to realize that the issue of lower-boundedness on the cubes could be resolved by introducing a double truncation. Most of the other details came down to writing out long and careful proofs of the measurability, showing all the indicated convergences actually occurred in the claimed spaces, and filling in the argument for showing higher integrability, realizing that the problem was really HLS in disguise.
This was a long and quite-classical paper, but I hope this rewriting has produced a useful and (more) legible map of this terrain, especially for those of us accustomed to more modern arguments for estimating singular integrals.
Work begun: August 31, 2025. Finished: September 2, 2025.
- Helpfully, some of the relevant calculations can be found (in great detail) in the excellent and highly-readable article by Ricardo G. Durán & Maria Amelia Muschietti, An Explicit Right Inverse of the Divergence Operator Which is Continuous in Weighted Norms, 148 Studia Math. 207 (2001), doi:10.4064/sm148-3-2. MR 1880723. (For more on this, and some interesting extensions, one may consult the helpful short guide by Gabriel Acosta & Ricardo G. Durán, Divergence Operator and Related Inequalities, SpringerBriefs in Mathematics, Springer, New York, 2017. MR 3618122.)
The remaining assertions listed above can be found, for instance, in section III.3 of G. P. Galdi, An Introduction to the Mathematical Theory of the Navier-Stokes Equations, 2d ed., Springer Monographs in Mathematics, Springer, New York, 2011. MR 2808162.
Both aforementioned books state generalizations of the domains to which this construction can be performed; the “John domains” are currently among the state-of-the-art; see Acosta & Durán. ↩︎ - Chérif Amrouche, On Some Equivalent Theorems: Poincaré, Korn, De Rham, Necas, Lions and Bogovskii. XIVth International Conference Zaragoza-Pau on Mathematics and its Applications (Sept. 12–15, 2016), available at https://hal.science/hal-02525742. ↩︎
- A. P. Calderón & A. Zygmund, On Singular Integrals, 78 Am. J. Math. 289 (1956), doi:10.2307/2372517. MR 0084633. ↩︎
- For more on this, including a discussion of known results and some directions of modern research, see Loukas Grafakos & Atanas Stefanov, \(L^p\) Bounds for Singular Integrals and Maximal Singular Integrals with Rough Kernels, 47 Ind. U. Math. J. 455 (1998), doi:10.1512/iumj.1998.47.1521. MR 1647912. ↩︎
- I.e., if we’re being truly honest, “best” here is really being used due to the unfortunate dearth of any competition whatsoever. ↩︎
- For instance, one inequality (important for bounding a certain kernel in \(L^q\)) was listed (rather unhelpfully!) as following from “a theorem of Young,” with a citation to the 1935 edition of Zygmund’s book, which I certainly did not have. Another (on the \(L \log L \to L^1_{\text{loc}}\) mapping properties of some singular integral operators) was stated with a citation to Calderón and Zygmund’s first paper for the argument; the original paper is rather irritating to read, and uses tools very far from the modern machinery and perspective. ↩︎
- A. P. Calderón & A. Zygmund, On the Existence of Certain Singular Integrals, 88 Acta Math. 85 (1952), doi:10.1007/BF02392130. MR 0052553. I apologize, but much of this paper (in style and in method) really is unfamiliar; rather unpleasant to struggle through. ↩︎
Leave a Reply