The T(1) Theorem, After David and Journé

Written

by

Update (October 7): The extension process for the bilinear form, from the \(L^2\)-dense subclass of \(\mathcal{S}\) to the whole Schwartz space, is nontrivial and was not checked or discussed in the original paper, and we did not prove this in the original write-up. Consequently, the original study guide should be considered incomplete, and the new one (PDF here; the old write-up is only left below for historical purposes) has been revised to address this issue. The second copy should be considered the canonical version. A post detailing the issues is here.

The write-up, as always, is here. More context and motivation is given below.


I decided that this summer (i.e., June to September, 2025) would be a good time to learn about \(T(1)\) theorems; as a consequence, I spent a few weeks building up to the original paper of David & Journé, which first kicked off the development of new testing conditions for investigating boundedness of singular integral operators.

I wrote a study guide for the paper, going through my usual process of taking every claim in the text and proving it in meticulous detail, including (especially) parts that had been discussed only briefly or with proofs omitted.

I chose to only review parts I–III of the article, focusing on the proof of the \(T(1)\) theorem; the other parts giving applications, though interesting, are still in the process of being read and may form the basis of future blog posts.

I indicate some of the major things I added, below:

  1. I include a few brief remarks, clarifying the nature of how we want the functional \(g \mapsto \langle T(1), g \rangle\) to act on \(C^{\infty}_c(\mathbb{R}^d) \cap H^1(\mathbb{R}^d)\), and how this is actually compatible with how we define action on \(L^{\infty}(\mathbb{R}^d)\) for genuine Calderón-Zygmund operators.
  2. In part II, I include the explicit details of the continuous paraproduct construction, carefully verifying measurability and continuity, smoothness, etc. when needed (as is my usual habit when reading these types of papers, in which such minor (but still important!) details are rarely verified).
  3. In part II, David and Journé assert that it is permissible to use a physical-space bump function to construct the paraproduct; they take a radial \(\varphi\) with unit mass and compact support in the ball \(B(0, 1)\). As it turns out, doing this would cause serious issues later down the line, when we need to verify a certain weakly-defined operator integral produces a difference of convolutions with approximate identities; we simply do not get smooth functions with such a choice. (This is essentially due to an issue with \(\xi \cdot (\nabla \widehat{\varphi})(\xi)\) vanishing at \(\xi = 0\), but not necessarily its derivatives; to avoid technicalities, we would need something like vanishing to infinite order.) So instead, we take the \(\varphi\) to be compactly-supported in frequency: \(\text{supp}(\widehat{\varphi}) \subseteq B(0, 1)\), with the Fourier transform radially-decreasing and constant in a neighborhood of the origin.1
  4. In the proof of Lemma 3 of the paper, we do have to be slightly careful, due to the Schwartz tails. We apply a bit more care in proving absolute integrability, and interchanging the orders of integrals, using some properties of \(\text{BMO}\) functions.
  5. For part III, we include a proof of the Cotlar-Stein lemma used in the paper, but not proven. This should make the entire reading experience self-contained. We also note that for part III, we actually can and must have bump functions compactly-supported in physical space; this is a key point of contrast between II and III, and perhaps the source of the confusion for why part II did the opposite of insisting on compact support in frequency-space.
  6. Finally, Lemma 4 of the paper includes a large number of properties supposed to be satisfied by the frequency-localized distributional kernels; most of these are only sketched briefly in the text, and not the hard ones (the ones that show the mean-zero conditions, for instance). Some of these (like the mean-zero ones) were formally-“obvious” to see, but required a bit of care with the proofs. Using some careful limiting arguments and a great deal of fidelity to the original definitions, we carefully verified each of these details.

The requirement of choosing frequency-localized (as opposed to physical-space localized) kernels, noted in item 3 above, also generates some other issues; for the most interesting one, there is a certain subtlety between (standard, \(L^1\)) integration against a \(\text{BMO}\) function, and the linear functional on \(H^1\) that it represents. We will use this to resolve a nagging question that might be very natural to those who have learned about duality on the \(H^1\).

The process of defining the action of a \(\text{BMO}\) function is famously subtle; Auscher, in his text on harmonic analysis, emphasizes this point (namely, that a general decomposition does not work if \(b \in \text{BMO}(\mathbb{R}^d) \setminus L^{\infty}(\mathbb{R}^d)\); see pp. 56–58 of the text).2 To briefly recall the details, we would first like to verify that for finite linear combinations of atoms, the operation of integration against a \(\text{BMO}\) function is a linear functional obeying an \(H^1\)-norm bound; to use the atomic decomposition, and get cancellation in each summand, one would like to interchange sum and integral. But this, unfortunately, is not permitted; convergence in \(H^1\) is very close to \(L^1\), and not generally strong enough to be paired with a \(\text{BMO}\) function; \(\text{BMO}\) is only exponentially-integrable, but not actually bounded.3

So to resolve this, we first truncate the \(\text{BMO}\) function in the integral to heights \(| \cdot | \leq k\); and interchange the sum in the atomic decomposition of a given element of the linear span of atoms. Only after we recover the norm bound do we consider the truncations; we send \(k \to \infty\) and use the regularity of \(\text{BMO}\) functions (viz., John-Nirenberg) to get convergence of the integral to the intended pairing, \(\langle b, f \rangle\) (as a genuine, absolutely-convergent integral).

From this, one gets a bound like $$|\langle b, f \rangle| \leq 2 \|b\|_{\text{BMO}(\mathbb{R}^d)} \|f\|_{H^1(\mathbb{R}^d)},$$ for all \(f\) in the linear span of atoms. This being a dense subspace, one then extends \(f \mapsto \langle b, f \rangle = \int_{\mathbb{R}^d} b \cdot f \, d\lambda^d\) to all of \(H^1(\mathbb{R}^d)\) by the usual density argument.

(This is, at least, one way to show it. Another way is to prove that part of the convergence in the atomic decomposition, for finite linear combinations of \(q\)-atoms in \(H^1(\mathbb{R}^d)\) (for any \(1 < q < \infty\)), is actually occurring in \(L^q(\mathbb{R}^d)\); as a consequence, by the regularity of \(b\), it will be permissible to interchange sum and integral.4 One should also note that corresponding results for the \(q = \infty\) endpoint (equating the infinite and finitary atomic decomposition norms for elements in the span of \(q = \infty\)-atoms) cannot be included, by a counterexample discussed by M. Bownik, which he attributes to Y. Meyer.5)

In any event, the procedure for defining \(\ell_b(f)\), where \(\ell_b\) is the bounded linear functional associated in this sense to the function \(b \in \text{BMO}(\mathbb{R}^d)\) (possibly unbounded),6 and \(f \in H^1(\mathbb{R}^d)\) is a general element of the Hardy space (without further regularity), is very much not obvious. The density procedure means that \(\ell_b(f) = \sum_{n = 1}^{\infty} \lambda_n \ell_b(a_n)\), using any atomic decomposition of \(f\); but what this means, concretely, is very unclear.

As part of this work, in order to deal with certain troublesome elements of the Hardy space (\(f \in \mathcal{S}(\mathbb{R}^d)\), with mean-zero but not necessarily of compact support), we proved the following result.

Intuitively, it gives that the evaluation of the linear functional coincides with our normal notion of duality pairing between functions, in the case when both possibilities make sense:

Theorem: Let \(f \in H^1(\mathbb{R}^d)\), \(b \in \text{BMO}(\mathbb{R}^d)\). Suppose \(f\) and \(b\) are such that the pointwise product \(b \cdot f\) lies in \(L^1(\mathbb{R}^d)\) (i.e., is absolutely-integrable). (For instance, any \(f \in \mathcal{S}(\mathbb{R}^d)\) with mean-zero will qualify.)

Then for the continuous linear functional \(\ell_b\) on \(H^1\) generated by \(b\), by the complicated extension process above, we have the equality $$\ell_b(f) = \int_{\mathbb{R}^d} b f \, d\lambda^d = \langle b, f \rangle.$$

This tells us, at least, that despite all the indirect and winding manipulations we went through to get to a continuous functional \(\ell_b\), for well-behaved (i.e., fast-decaying) functions, its behavior really doesn’t depart from our understanding.7

Work begun: September 5, 2025. Finished: September 6, 2025.

Revised: October 6, 2025 (see new post).


  1. My general strategy for writing these things is to first read through the article, working out all the calculations and on separate pieces of paper until I’m sure that I’ve gotten all the details nailed down, and then proceed to begin writing out the note. In this case, my family had just returned from a trip to Arizona, Nevada, and Utah; I had brought the paper with me and had been working on this calculation on the plane. So, as we were descending into SJC, I was scribbling all sorts of notes and calculations across pieces of paper, while the passenger in the adjacent seat watched me work. In retrospect, I’m sure she had reason to think I was crazy. In any event, that was the moment I realized the choice of approximate identity for the continuous paraproduct had to be changed. ↩︎
  2. See Pascal Auscher (with Lashi Bandara), Real Harmonic Analysis, 3d edition, ANU Press, Canberra, Australia, 2022 (revised; 2d ed., 2016; 1st ed., 2012). doi:10.22459/RHA.2012. ↩︎
  3. If one is very motivated by the technical details, one could link this to the dual \(L \log L\) Orcliz-type space, and observe that while the maximal function does give such a bound (which is, in a sense, “equivalent,” as we now know due to Stein), the grand maximal function (by not having the absolute value inside) does not really allow for \(L \log L\) control, due to cancellation. Another way to say this is that we have \(\mathcal{M} f \lesssim M f\), but not the other way around. In general, by “building up” a singularity at a point, one can create an \(H^1\) function that is only in \(L^1\), not in \(L \log L\). This also has implications for the result that \(f \in H^1(\mathbb{R}^d)\) will obey \(f \in (L \log L)_{\text{loc}}(U)\) for all open \(U \subseteq \{f > 0\}\), but that this cannot be strengthened to remove the separation restriction \(_{\text{loc}}\). Thanks to Professor Killip for pointing this out in office hours. ↩︎
  4. This is a result due to Stefano Meda, Peter Sjögren & Maria Vallarino, On the \(H^1\)-\(L^1\) Boundedness of Operators, 136 Proc. Am. Math. Soc. 2921 (2008). doi:10.1090/S0002-9939-08-09365-9. MR 2399059.
    Despite Auscher (and the MR review of the article, by Bownik) proclaiming this argument to be “intricate,” or “subtle” (resp.), this really isn’t too complicated, at least in my mind: one is essentially repeating the stratification at each \(k \in \mathbb{Z}\), and in the high components, using the unnormalized atoms to form building blocks for a layer at height \(\sim 2^k < \mathcal{M} f\). With the trick the authors use (of expanding the distances in the Whitney decomposition), each atom really is like a small characteristic function, and they form a finitely-overlapping covering of the entire \(\{\mathcal{M} f > 2^k\}\) set. Mentally, I find it easy to see the layers being built, one on top of the other, stopping at each point once it leaves the open sets (and thus, by summing the geometric series, the tower reaches a comparable height to the maximal function). The atomic decomposition is literally working under the hood of \(\mathcal{M} f\). So it is a consequence that the sum with absolute values is pointwise dominated by the grand maximal function, and by the radial majorant lemma, depending on the \(N\) and \(A\) parameters, this is pointwise dominated by \(\leq C(d, N, A) \cdot M f\). Thus we have dominated convergence in \(L^q\). (Maybe this familiarity is a reflection of how much time — perhaps too much? lol — that I’ve spent, thinking about the details of the proof of the atomic decomposition. To be honest, I’ve been rereading (and rewriting) that section of my notes ever since the argument was presented in 247B last winter.) ↩︎
  5. Marcin Bownik, Boundedness of Operators on Hardy Spaces via Atomic Decompositions, 133 Proc. Am. Math. Soc. 3535 (2005). doi:10.1090/S0002-9939-05-07892-5. MR 2163588. ↩︎
  6. We can, at least, clearly see that for any two continuous linear functionals both associated with a function like this (and thus agreeing on atoms), the two will coincide on \(H^1\) by continuity and density. ↩︎
  7. This fact, of the intuition-aligning type, was somewhat involved to prove, and required taking limits in the Banach-Alaoglu sense, and verifying that the convergent subnets actually all tended to the same thing (i.e., the functional \(\ell_b\); essentially by the agreement-on-atoms-implies-agreement-for-all principle discussed in footnote 5, above), then considering how they converged on the physical-space side. I had to check Folland, Ch. 4 & 5, many times when writing this. In retrospect, this could have been easier if I had shown first that the predual \(H^1(\mathbb{R}^d)\) was separable, so that we could use the straightforward subsequential form of Banach-Alaoglu; sequences, of course, allow for dominated convergence. ↩︎

Leave a Reply

(In)Complete Thoughts

Mathematical musings & miscellany

Discover more from (In)Complete Thoughts

Subscribe now to keep reading and get access to the full archive.

Continue reading