Our First Correction!

Written

by

Today I’ve decided that I should seriously revise a past write-up (and post). I did not independently verify a density/extension process from an argument, even though it had been omitted from the original paper; this post is my argument to close that gap for good.

The new and corrected write-up is here. The old post, here, has been updated to link to this current one.

The \(T(1)\) theorem of David & Journé (1984) is a landmark result of harmonic analysis,1 and in greatest generality, precisely characterizes the situation where a distribution-valued operator \(T : \mathcal{S}(\mathbb{R}^d) \to \mathcal{S}'(\mathbb{R}^d)\) is given by a Calderón-Zygmund operator.

The general \(T(1)\) result of David & Journé (unlike, say, Tao or Tolsa’s formulations of the \(T(1)\) theorem) works with distributions on Schwartz functions; this seriously restricts the functions on which we can test our operator (for instance, no characteristic functions of balls or cubes anymore), and means we cannot ask about pointwise values on the support of the input; everything must be processed from a distance, through a dual pairing.

This introduces more technicalities, but it also produces a stronger result: rather than assuming the operator is generated by an integral kernel that is assumed bounded and slice-integrable (in the sense of Tao or Tolsa), we simply work with a distribution-valued operator on the Schwartz space, which has much smaller natural domain and could a priori be much rougher.


We briefly outline the proof of the general \(T(1)\) theorem, and point out the difficulties where they arise.

The first step is to take the original distribution-valued operator \(T\) to induce three specific sequences of operators \((T_j)\), \((T_j’)\), \((T_j^{ \prime \prime})\) indexed over \(\mathbb{Z}\), each of which can be found to act via genuine integral-kernel operators on Schwartz functions. Some estimates of their kernels show that they are summable in the sense of the Cotlar-Stein almost-orthogonality lemma.

So we may consider the operator $$\tilde{T} = \sum_j T_j + T_j’ \ – \, T_j^{\prime \prime},$$ which converges unconditionally in the strong operator topology. Also, by some direct calculations using the precise definitions of the various operators, we have a telescoping-sum type result $$\sum_{|j| \leq M} \langle (T_j + T_j’ \ – \, T_j^{\prime \prime})(f), g \rangle = \langle T(P_{\leq – M} f), P_{\leq – M} g \rangle \ – \, \langle T(P_{\leq M + 1} f), P_{\leq M + 1} g \rangle,$$ where the \(P_M\) here are not the Littlewood-Paley frequency localization operators, but instead convolution with \(2^{- M d} \varphi( \cdot / 2^M)\), \(\varphi\) a compactly-supported bump function of unit mass. So the operator \(P_{\leq – M}\), for \(M\) large and positive, will act as an approximate identity; \(P_{\leq M + 1}\) acts as an averaging operator against an extremely-low-frequency function, almost constant as \(M\) grows large.

To obtain their result, David and Journé work with a certain class of functions: the \(f, g \in \mathcal{S}(\mathbb{R}^d)\) with \(\widehat{f}\) vanishing in a neighborhood of the origin, \(\xi = 0\). This is because for such \(f\) and \(g\), $$P_M f \to f$$ in the topology of \(\mathcal{S}(\mathbb{R}^d)\). Since convolutions with approximate identities naturally converge in the topology of the Schwartz space, it would follow that the \(M \to \infty\) limit would produce $$\langle T(f), g \rangle = \langle \tilde{T}(f), g \rangle.$$

It is important to note that for general \(f \in \mathcal{S}(\mathbb{R}^d)\), it is very much untrue that \(P_{\leq M} f\) converges to \(0\) in \(\mathcal{S}(\mathbb{R}^d)\); it is a special feature of this class that such vanishing occurs, which at least partially justifies the decision to consider it, I suppose.

At this point, after verifying that the Cotlar-Stein lemma can be applied, the proof in David & Journé (1984) ends, and the focus moves on to discussing applications.

This is very unsatisfying, because while \(\tilde{T}\) is an \(L^2\)-bounded operator, and Schwartz functions with compact frequency support away from the origin or infinity are dense in \(L^2\), they are emphatically not dense in the Schwartz space. Indeed, their closure should be something like the functions vanishing to infinite order at the frequency origin.

Consequently, the fact that the identity \(T = \tilde{T}\) holds for this special subclass of Schwartz functions (we will abbreviate this as \((\mathcal{S})_0\)) places us in a very confusing situation. It is then true that the (very) restricted distribution-valued operator \(T \mid_{(\mathcal{S})_0} : (\mathcal{S})_0 \to ((\mathcal{S})_0)’\) obeys \(L^2\) bounds; this has a unique extension to all of \(L^2\), by density, in the standard way; but it is very unclear whether this extension is actually the original operator \(T\).

This is the source of the difficulty; any extension beyond this non-dense subclass is very difficult to do in \(\mathcal{S}(\mathbb{R}^d) \times \mathcal{S}(\mathbb{R}^d)\), since the distance between an element of that class and any function with nonzero mean is very, very far. At least on the level of Schwartz kernels and tempered distributions, we cannot bridge this gap.

(This difficulty does not appear in the Tao or Tolsa proofs of the \(T(1)\) theorems, for instance, because their qualitative hypotheses on \(T\) ensured that it was actually given by a qualitatively \(L^2\)-bounded operator; the focus of the theorem was to derive quantitative norm bounds; consequently, taking \(L^2\) limits would preserve both sides of the equality of pairings, whereas \(L^2\) convergence of elements of \((\mathcal{S})_0\) to an element of \(\mathcal{S}\) is far too rough to suffice when we are applying tempered distributions and Schwartz seminorms.)

It was this difficulty that consumed me for the past few days; originally, after finishing the \(T(1)\) theorem paper, I did not see any difficulties with the extension argument, perhaps because I had read the Tao and Tolsa proofs first and was still thinking in \(L^2\). However, the difficulty of the extension in Schwartz space gradually impressed itself on me, and I finally cleared some time to think about this problem yesterday and worked it out.

So we may summarize the task before us:

  • Verify that the operator produced by the sum \(\sum_j T_j + T_j’ \ – \, T_j^{\prime \prime}\) is actually equal to the operator \(T\), on the entire Schwartz space.

To do this, we must actually rewrite the proof of the \(T(1)\) theorem somewhat. The key is to take a dual pairing $$\left \langle \sum_{|j| \leq M} (T_j + T_j’ \ – \, T_j^{\prime \prime})(f), g \right \rangle = \langle T(P_{\leq – M} f), P_{\leq – M} g \rangle \ – \, \langle T(P_{\leq M + 1} f), P_{\leq M + 1} g \rangle$$ for arbitrary \(f, g \in C^{\infty}_c(\mathbb{R}^d)\), rather than taking elements of \((\mathcal{S})_0\).2 Indeed, we will not use anything from \((\mathcal{S})_0\) here initially, as it is vital (for our next application) for the relevant objects to have compact support.

We will then see that the low-frequency term (the second term here, being subtracted) vanishes in the \(M \to \infty\) limit; this will not be because the \(P_{\leq M} f\) vanish in \(\mathcal{S}(\mathbb{R}^d)\) as \(M\) tends to infinity, which would be what would be needed if we were trying to close the argument with the continuity of \(T : \mathcal{S}(\mathbb{R}^d) \to \mathcal{S}'(\mathbb{R}^d)\); instead, we will use another, more specific hypothesis:

  • Weak Boundedness Property: Let \(\mathcal{B} \subseteq C^{\infty}_c(\mathbb{R}^d)\) be a bounded family (in the sense of the Definition in David & Journé; if one wishes to be concrete, sometimes this is stated in terms of the slightly-more-stringent (but definitionally-cleaner) notion of “normalized bumps”). Then there exists a uniform constant such that $$|\langle T(\phi^{x, R}), \psi^{x, R} \rangle| \leq C_{\mathcal{B}} R^d$$ for all \(\phi, \psi \in \mathcal{B}\), where we adopt the abbreviation \(f^{x, R} = f( (\cdot \ – \, x) / R)\), for \(x \in \mathbb{R}^d\), \(R > 0\).

Now, in the argument given in David & Journé, we mainly use the weak-boundedness property to control the nearby-interactions (\(x \approx y\)) in the integral kernels \(K_j\), \(K_j’\), \(K_j^{\prime \prime}\). It is interesting to note that this property also comes in at the end, to remove the final obstruction to obtaining the representation for \(T\) on all Schwartz functions.

Roughly-speaking, one writes the final error term \(\langle T(P_{\leq M + 1} f), P_{\leq M + 1} g \rangle\) as an integral operator on \(f\) and \(g\), with a kernel of \(K_M(x, y) = \langle T(\varphi_{y, M}), \varphi_{x, M} \rangle\) for various functions \(\varphi\), suitably-translated and dilated depending on all the parameters, and one uses the weak-boundedness property to show that this vanishes (!) as \(M \to \infty\).

From this, one obtains the desired result, $$\langle T(f), g \rangle = \langle \tilde{T}(f), g \rangle,$$ for all \(f, g \in C^{\infty}_c(\mathbb{R}^d)\); it is then straightforward to extend this to all Schwartz functions, using the continuity of \(T\).

And from this, we have the bound on the dual pairing of $$|\langle T(f), g \rangle| \lesssim \|f\|_{L^2(\mathbb{R}^d)} \|g\|_{L^2(\mathbb{R}^d)},$$ valid for all \(f, g \in \mathcal{S}(\mathbb{R}^d)\); this allows us to conclude that each \(T(f)\), for any \(f \in \mathcal{S}(\mathbb{R}^d)\) (not just \(f \in (\mathcal{S})_0\)), is actually given by a \(L^2(\mathbb{R}^d)\) function (not just an \(L^2\) function plus a polynomial, which would be the strongest we could conclude if we were only testing the distribution against functions \(g \in (\mathcal{S})_0\)).


Finally, we should note that in the process of examining this proof, and going back to its fine details, I found that it seems this same \(T(1)\) theorem holds for operators \(T : \mathcal{D} \to \mathcal{D}’\), not even requiring \(T : \mathcal{S} \to \mathcal{S}’\) (the latter is strictly stronger, as the domain is larger while the range is smaller); one does not really need to test on Schwartz functions for this argument. The \(\text{BMO}\) conditions on \(T(1)\), \(T^t(1)\) are verified by testing on elements of \(H^1(\mathbb{R}^d) \cap C^{\infty}_c(\mathbb{R}^d)\), while the construction of the almost-orthogonal sequences of operators only involves evaluating the operator \(T\) on pairs of compactly-supported functions, translated and dilated according to various parameters. The only times where we used that \(T\) can be evaluated on elements of \(\mathcal{S}\) was in the very end, when we extended \(\langle T(f), g \rangle\) from \(C^{\infty}_c(\mathbb{R}^d) \times C^{\infty}_c(\mathbb{R}^d)\) to \(\mathcal{S}(\mathbb{R}^d) \times \mathcal{S}(\mathbb{R}^d)\). This isn’t really some novel insight, but it would seem to widen the applicability of the result, slightly.


  1. I’m unsure whether the \(T(1)\) theorem should be described as belonging to the field of “classical” or “modern” harmonic analysis, as I realized when I was typing the above sentence. This is probably an indication that I’m not yet familiar-enough with my field. ↩︎
  2. This is also more in line with part III of David & Journé, which largely works with physical-space kernels and does little on the frequency side; the use of frequency-side conditions in physical-space contexts, or vice versa, can introduce technicalities that must later be resolved, as we discussed in the previous (original) post on David and Journé’s paper regarding their construction of the continuous paraproduct. ↩︎

Leave a Reply

(In)Complete Thoughts

Mathematical musings & miscellany

Discover more from (In)Complete Thoughts

Subscribe now to keep reading and get access to the full archive.

Continue reading